← All Data Science Flashcard Decks

Data Cleaning and Preparation Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Data Cleaning and Preparation flashcards as text
  1. The interquartile range (IQR) method is commonly used to detect:

    Answer: Outliers

    Values outside 1.5×IQR beyond the quartiles are flagged as outliers.

  2. Min-max scaling transforms a feature to which range by default?

    Answer: [0, 1]

    Min-max scaling rescales values to the 0 to 1 range.

  3. Z-score standardization produces data with:

    Answer: Mean of 0 and standard deviation of 1

    Standardization centers data to mean 0 and unit standard deviation.

  4. Which encoding creates a separate binary column for each category value?

    Answer: One-hot encoding

    One-hot encoding generates one binary indicator column per category.

  5. Why scale features before training a k-nearest-neighbors model?

    Answer: Distance calculations are sensitive to feature magnitude

    KNN uses distances, so unscaled large-magnitude features dominate.

  6. Applying a log transform to a right-skewed feature primarily aims to:

    Answer: Reduce skew and compress large values

    Log transforms compress large values and reduce right skew.

  7. Removing a data point with age = 250 from a human dataset is best described as:

    Answer: Handling an invalid/implausible value

    An impossible age is an invalid value to correct or remove.