โ† All Data Science Flashcard Decks

FREE Data Science Feature Engineering and Selection Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 FREE Data Science Feature Engineering and Selection Questions and Answers flashcards as text
  1. Which dimensionality reduction technique is best suited for preserving local neighborhood structure in high-dimensional data?

    Answer: t-SNE

    t-SNE preserves local pairwise distances, making it effective for visualizing clusters in high-dimensional data.

  2. What is the purpose of applying a Box-Cox transformation to a feature?

    Answer: To make the feature distribution more closely approximate a normal distribution

    Box-Cox transformation applies a power function to stabilize variance and make data more Gaussian-like.

  3. When using L1 regularization (Lasso) for feature selection, what happens to irrelevant feature coefficients?

    Answer: They are driven exactly to zero

    L1 regularization penalizes the absolute value of coefficients, forcing unimportant ones to exactly zero.

  4. What is feature hashing (the hashing trick) primarily used for?

    Answer: Reducing high-cardinality categorical features to a fixed-size vector

    Feature hashing maps categorical values to a fixed number of buckets using a hash function, handling high cardinality efficiently.

  5. Which metric is commonly used in filter-based feature selection to rank features for a classification task?

    Answer: Chi-squared statistic

    The chi-squared test measures the dependence between a categorical feature and the target class, making it ideal for filter-based ranking.

  6. What problem does feature scaling solve when using K-Nearest Neighbors for prediction?

    Answer: Prevents features with larger numeric ranges from dominating the distance calculation

    Feature scaling ensures all features contribute equally to distance metrics by normalizing their ranges.