Data Wrangling and Preprocessing Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Wrangling and Preprocessing flashcards as text
A column of customer ages contains the value 999 for records where age was unknown. What is this an example of?
Answer: A sentinel value encoding missing data
Sentinel values like 999 are placeholders that secretly encode missing or unknown data and should be converted to NaN before analysis.
In pandas, which method removes rows containing any missing values from a DataFrame?
Answer: df.dropna()
dropna() drops rows (or columns) that contain NaN values by default.
You merge two tables on a customer ID, but some IDs exist only in the left table. Which join keeps all left-table rows?
Answer: Left join
A left join retains every row from the left table and fills unmatched right-table columns with nulls.
A date column is stored as strings like '2026-06-30'. What preprocessing step lets you extract the month easily?
Answer: Convert the column to a datetime type
Parsing strings into a datetime type exposes attributes like .month, .year, and .dayofweek.
Which technique replaces missing numeric values with the column's median?
Answer: Median imputation
Median imputation fills NaNs with the median, which is robust to outliers compared to the mean.
Why is the median often preferred over the mean for imputing a skewed income column?
Answer: It is less affected by extreme high values
The median is resistant to outliers, so skewed distributions don't distort the imputed value as much as the mean would.
What does 'tidy data' mean in the context of data wrangling?
Answer: Each variable is a column and each observation is a row
Tidy data has one variable per column, one observation per row, and one observational unit per table.