Exploratory Data Analysis Techniques Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Exploratory Data Analysis Techniques flashcards as text
Which plot is most appropriate for visualizing the joint distribution of two continuous variables along with their marginal distributions?
Answer: Joint plot (with marginal histograms)
A joint plot displays a scatter or density plot in the center with marginal histograms or KDE curves along each axis, showing both joint and individual distributions.
You notice that a numeric column has values ranging from 1 to 1,000,000 with most values under 1,000. Which transformation would best reveal structure in the lower range?
Answer: Logarithmic transformation
A log transformation compresses large values while expanding small values, making structure in the dense lower range visible without losing the high-end variation.
What does a high kurtosis value (leptokurtic distribution) indicate about a dataset?
Answer: The distribution has heavy tails and a sharp peak
Leptokurtic distributions have excess kurtosis greater than 0, indicating heavier tails and more frequent extreme values than a normal distribution.
Which EDA technique would you use to detect multicollinearity among predictor variables before modeling?
Answer: Correlation matrix or variance inflation factors
A correlation matrix reveals linear dependencies between predictors, and VIF quantifies how much one predictor's variance is explained by others.
In EDA, what is a 'rug plot'?
Answer: Tick marks along an axis showing individual data point locations
A rug plot places short vertical tick marks along an axis at each data point's position, showing the actual distribution of individual observations.
When should you prefer a log scale on a histogram's x-axis?
Answer: When the variable spans several orders of magnitude
A log scale is ideal for variables spanning several orders of magnitude (e.g., income, population) so that all ranges are visually represented proportionally.
What does the Interquartile Range (IQR) measure?
Answer: The spread of the middle 50% of the data
IQR = Q3 − Q1 and captures the spread of the central half of the data, making it robust to extreme outliers.