โ† All Data Science Flashcard Decks

FREE Data Science Data Wrangling and Preprocessing Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 FREE Data Science Data Wrangling and Preprocessing Questions and Answers flashcards as text
  1. What is the primary risk of using Min-Max scaling on a dataset before splitting into training and test sets?

    Answer: It introduces data leakage because test set statistics influence the transformation

    Fitting the scaler on the entire dataset before splitting causes data leakage because the minimum and maximum values from the test set influence the training data transformation.

  2. Which pandas method is most efficient for converting a wide-format DataFrame with repeated measurements into long format?

    Answer: melt()

    The melt() function unpivots a wide-format DataFrame into long format by converting column headers into row values under a new variable column.

  3. When dealing with a heavily right-skewed target variable in a regression task, which preprocessing step is most commonly applied?

    Answer: Box-Cox or log transformation

    Box-Cox and log transformations compress large values and spread small values, making the distribution more symmetric and helping linear models meet their normality assumptions.

  4. What is the purpose of using stratified sampling when splitting a dataset into training and test sets?

    Answer: To maintain the same class distribution in both the training and test sets

    Stratified sampling preserves the proportion of each class in both splits, which is critical for imbalanced datasets where random splitting could leave minority classes underrepresented in one set.

  5. Which approach is recommended for handling high-cardinality categorical features with hundreds of unique values?

    Answer: Target encoding or frequency encoding

    Target encoding replaces categories with the mean of the target variable for that category, while frequency encoding uses occurrence counts, both avoiding the dimensionality explosion of one-hot encoding.

  6. What does the Interquartile Range (IQR) method define as an outlier?

    Answer: Any value below Q1 minus 1.5 times IQR or above Q3 plus 1.5 times IQR

    The IQR method flags values below Q1 - 1.5*IQR or above Q3 + 1.5*IQR as outliers, which is robust to non-normal distributions unlike standard deviation methods.