Data Cleaning and Preparation Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Cleaning and Preparation flashcards as text
Combining two DataFrames on a shared key column in pandas is done with:
Answer: pd.merge()
pd.merge joins DataFrames on common key columns.
An 'inner join' between two tables returns:
Answer: Only rows with matching keys in both tables
An inner join keeps only keys present in both tables.
MICE (Multiple Imputation by Chained Equations) is a method for:
Answer: Imputing missing values using other features
MICE models each feature with missing data from the others to impute values.
Reshaping data from wide to long format in pandas typically uses:
Answer: pd.melt()
pd.melt unpivots wide columns into long key-value rows.
When categorical labels have a natural order (low, medium, high), the best encoding is:
Answer: Ordinal encoding
Ordinal encoding preserves the inherent rank order of categories.
A 'data dictionary' in a preparation workflow primarily documents:
Answer: Each field's meaning, type, and allowed values
A data dictionary defines column meanings, types, and valid values.
SMOTE is applied during preparation mainly to address:
Answer: Class imbalance by synthesizing minority-class samples
SMOTE oversamples the minority class by creating synthetic examples.