โ† All AML Flashcard Decks

AML Feature Engineering & Data Preprocessing Flashcards

6 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 AML Feature Engineering & Data Preprocessing flashcards as text
  1. What is the key difference between normalization and standardization?

    Answer: Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance

    Normalization maps values to a fixed range like [0,1] while standardization transforms data to have mean 0 and standard deviation 1.

  2. Which feature engineering technique creates new features from products or powers of existing features?

    Answer: Polynomial feature expansion

    Polynomial feature expansion generates new features as products or powers of existing ones, enabling linear models to capture non-linear patterns.

  3. What is target encoding for categorical variables?

    Answer: Replacing each category value with the mean of the target variable for that category

    Target encoding substitutes each category with the average target value for that category, directly capturing predictive power.

  4. Why is k-fold cross-validation used when evaluating feature engineering choices?

    Answer: To obtain a more reliable performance estimate across multiple data subsets

    K-fold cross-validation rotates the validation set across k subsets, giving a more robust estimate than a single train-test split.

  5. Which method identifies predictive features by measuring how much each reduces impurity in a trained Random Forest?

    Answer: Feature importance from Random Forest

    Random Forest models compute feature importance scores based on average impurity reduction contributed by each feature across all trees.

  6. What does the Variance Inflation Factor (VIF) measure in regression analysis?

    Answer: The degree of multicollinearity among predictor variables

    VIF quantifies how much the variance of a regression coefficient is inflated due to correlation with other predictor variables.