AML Feature Engineering & Data Preprocessing Flashcards
6 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 AML Feature Engineering & Data Preprocessing flashcards as text
What is the key difference between normalization and standardization?
Answer: Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance
Normalization maps values to a fixed range like [0,1] while standardization transforms data to have mean 0 and standard deviation 1.
Which feature engineering technique creates new features from products or powers of existing features?
Answer: Polynomial feature expansion
Polynomial feature expansion generates new features as products or powers of existing ones, enabling linear models to capture non-linear patterns.
What is target encoding for categorical variables?
Answer: Replacing each category value with the mean of the target variable for that category
Target encoding substitutes each category with the average target value for that category, directly capturing predictive power.
Why is k-fold cross-validation used when evaluating feature engineering choices?
Answer: To obtain a more reliable performance estimate across multiple data subsets
K-fold cross-validation rotates the validation set across k subsets, giving a more robust estimate than a single train-test split.
Which method identifies predictive features by measuring how much each reduces impurity in a trained Random Forest?
Answer: Feature importance from Random Forest
Random Forest models compute feature importance scores based on average impurity reduction contributed by each feature across all trees.
What does the Variance Inflation Factor (VIF) measure in regression analysis?
Answer: The degree of multicollinearity among predictor variables
VIF quantifies how much the variance of a regression coefficient is inflated due to correlation with other predictor variables.