← All MS-DS Master of Data science Flashcard Decks

Exploratory Data Analysis Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Exploratory Data Analysis flashcards as text
  1. During EDA you discover that a numeric feature has a bimodal distribution. What does this most likely suggest?

    Answer: There may be two distinct subpopulations in the data

    Two peaks in a distribution often indicate the presence of two underlying groups or processes that were combined in the dataset.

  2. Which dimensionality reduction technique is commonly used in EDA to project high-dimensional data onto 2D for visual cluster discovery?

    Answer: t-SNE

    t-SNE (t-distributed Stochastic Neighbor Embedding) is designed to preserve local neighborhood structure, making clusters visible in 2D projections.

  3. When examining a scatter plot of two variables, you observe a curved (non-linear) pattern rather than a straight line. Which correlation measure is MORE appropriate than Pearson's r?

    Answer: Spearman's rank correlation

    Spearman's correlation measures monotonic association using ranks rather than raw values, handling non-linear but monotonic relationships better than Pearson's r.

  4. A feature has the values: [1, 2, 3, 4, 100]. The z-score of 100 is approximately 2.18. Should 100 be flagged as an outlier using the z-score method with threshold |z| > 3?

    Answer: No, because its z-score does not exceed 3

    The z-score method only flags values with |z| > 3 as outliers; z ≈ 2.18 is below that threshold, so 100 would not be flagged by this rule alone.

  5. Which EDA technique is specifically designed to reveal how the marginal distribution of one variable changes across levels of a second categorical variable?

    Answer: Box plots grouped by category

    Side-by-side box plots, one per category level, directly compare the center, spread, and skew of the numeric variable across groups.

  6. What is the primary risk of using only summary statistics (mean, std) without visualizing data during EDA?

    Answer: Different datasets can share identical summary statistics while having very different distributions (Anscombe's quartet)

    Anscombe's quartet demonstrates that four datasets with nearly identical means, variances, and correlations can have completely different visual patterns and structures.

  7. In EDA, what does a 'long tail' in a distribution imply for downstream machine learning modeling?

    Answer: Rare extreme values may disproportionately influence models sensitive to scale or magnitude

    Long-tailed features can distort distance-based or linear models unless transformed, requiring EDA to flag them for preprocessing decisions.