Feature Engineering and Selection Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Feature Engineering and Selection flashcards as text
Tree-based models provide feature importance primarily based on:
Answer: How much each feature reduces impurity across splits
Impurity-based importance measures each feature's total reduction in impurity over the tree splits.
Permutation importance measures a feature's value by:
Answer: Shuffling its values and observing the drop in model performance
Permutation importance randomly shuffles a feature and records how much performance degrades.
Which dimensionality reduction technique is unsupervised and linear?
Answer: Principal Component Analysis (PCA)
PCA is an unsupervised, linear method that projects data onto directions of maximum variance.
A key difference between PCA and feature selection is that PCA:
Answer: Creates new combined features rather than keeping original ones
PCA produces transformed components, whereas selection retains a subset of original features.
For text data, TF-IDF weighting down-weights words that:
Answer: Appear frequently across many documents
TF-IDF reduces the weight of common words appearing in many documents, emphasizing distinctive terms.
Feature hashing (the hashing trick) is mainly used to:
Answer: Map high-cardinality categories to a fixed-size space efficiently
Feature hashing maps many categories into a fixed number of buckets, saving memory at the cost of collisions.
When using SHAP values for feature analysis, they explain:
Answer: Each feature's contribution to an individual prediction
SHAP values attribute a prediction's deviation from baseline to each contributing feature.