Data Wrangling and Preprocessing Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Wrangling and Preprocessing flashcards as text
When concatenating training feature engineering, fitting an encoder on the full dataset before train/test split causes what problem?
Answer: Data leakage
Fitting any preprocessing on the combined data lets test information influence training, which is data leakage.
A target variable has 95% class A and 5% class B. Which preprocessing technique addresses this imbalance?
Answer: Resampling such as SMOTE or undersampling
Oversampling the minority class (e.g., SMOTE) or undersampling the majority class rebalances skewed class distributions.
Why should imputation of missing values inside a cross-validation loop be done within each fold rather than once beforehand?
Answer: To avoid leaking information across folds
Imputing before splitting lets statistics from validation folds leak into training, inflating performance estimates.
Two highly correlated numeric features (r=0.98) are kept in a linear model. What issue does this raise?
Answer: Multicollinearity
Highly correlated predictors cause multicollinearity, making coefficient estimates unstable and hard to interpret.
A free-text 'notes' column needs to become model features. Which preprocessing step is most appropriate first?
Answer: Text vectorization (e.g., TF-IDF or tokenization)
Raw text must be tokenized and vectorized (e.g., TF-IDF) into numeric features before most models can use it.
When using scikit-learn, bundling scaling, encoding, and the model into a single object that fits in order prevents leakage and simplifies deployment. What is it called?
Answer: A Pipeline
A Pipeline chains preprocessing steps and the estimator so transforms are fit only on training data within each fold.
You scale numeric columns but leave one-hot encoded binary columns unscaled. Why is this generally acceptable?
Answer: Binary 0/1 columns are already on a comparable bounded scale
One-hot columns are already bounded to 0 and 1, so they are on a scale comparable to standardized numeric features.