Data Cleaning and Preparation Flashcards
6 cards from real DA practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Cleaning and Preparation flashcards as text
What is the most common method for handling missing numeric values in a dataset?
Answer: Replace with the column mean, median, or mode
Imputing missing values with the mean, median, or mode preserves dataset size and is a standard starting point for handling missingness.
What is data normalization?
Answer: Scaling numeric values to a common range such as 0 to 1
Normalization rescales features to a standard range so that no single feature dominates due to its magnitude.
Which of the following best describes an outlier?
Answer: A value significantly different from the rest of the dataset
An outlier is a data point that lies far outside the typical range of values and can skew analysis results.
What does data deduplication mean?
Answer: Identifying and removing duplicate records
Deduplication removes redundant rows that represent the same entity, preventing double-counting in analysis.
What is a data type mismatch issue?
Answer: When a column stores values in an incorrect format for its intended data type
A data type mismatch occurs when values are stored in the wrong format, such as dates stored as strings, causing calculation errors.
What is one-hot encoding used for in data preparation?
Answer: Converting categorical variables into binary numeric columns
One-hot encoding transforms each category level into a separate binary column (0 or 1), enabling algorithms that require numeric input to process categorical data.