โ† All MS-DS Master of Data science Flashcard Decks

Feature Engineering Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Feature Engineering flashcards as text
  1. Which technique is used to reduce the impact of outliers when scaling numerical features?

    Answer: Robust Scaling

    Robust Scaling uses the median and interquartile range (IQR) instead of mean and standard deviation, making it resistant to the influence of outliers.

  2. What is the primary purpose of one-hot encoding in feature engineering?

    Answer: Convert nominal categorical variables into binary indicator columns

    One-hot encoding converts nominal categorical variables into binary (0/1) indicator columns so that machine learning algorithms can process them without implying an ordinal relationship.

  3. Label encoding is most appropriate for which type of categorical variable?

    Answer: Ordinal variables with a meaningful rank order

    Label encoding assigns integer values that preserve rank order, making it appropriate for ordinal variables where the numerical order is meaningful (e.g., low=1, medium=2, high=3).

  4. Which feature transformation is most commonly applied to correct a right-skewed (positively skewed) distribution?

    Answer: Applying a log transformation

    A log transformation compresses large values more than small values, which reduces right skew and helps bring the distribution closer to normal.

  5. What does 'feature interaction' refer to in the context of feature engineering?

    Answer: Creating new features by combining two or more existing features

    Feature interaction involves creating new features by combining existing ones (e.g., multiplication or ratio), which can capture relationships that individual features cannot represent.

  6. Which imputation strategy is most appropriate when data is missing completely at random (MCAR) and the dataset is small?

    Answer: Multiple imputation using chained equations (MICE)

    MICE (Multiple Imputation by Chained Equations) is preferred for small datasets with MCAR missingness because it preserves variability and uses information from other features to impute realistically.

  7. What is target encoding and what is its main risk?

    Answer: Replacing a category with the mean of the target variable for that category; risk is data leakage

    Target encoding replaces each category with the mean target value for that category, which is powerful but risks leaking target information into features if not properly cross-validated.

Feature Engineering Flashcards โ€” MS-DS Master of Data science Study Cards with Answers