Statistical Inference & Regression Models Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Statistical Inference & Regression Models flashcards as text
In multiple regression, the partial F-test is used to:
Answer: Compare a reduced model to a full model by testing a subset of predictors simultaneously
The partial F-test evaluates whether a group of predictors collectively adds significant explanatory power beyond what the reduced model already explains.
A Q-Q plot of regression residuals that shows heavy tails indicates:
Answer: Departure from normality with more extreme values than a normal distribution
Heavy tails in a Q-Q plot mean the residuals have a leptokurtic distribution with more extreme outliers than expected under normality.
Which scenario correctly describes an interaction term in multiple regression?
Answer: The effect of X₁ on Y depends on the value of X₂
An interaction term (X₁ × X₂) captures the situation where the marginal effect of one predictor on the response changes depending on another predictor's value.
The bootstrap method is used in regression to:
Answer: Estimate the sampling distribution of an estimator by resampling with replacement
Bootstrap resampling creates many pseudo-samples from the observed data to empirically estimate the sampling distribution and standard errors of estimators.
In generalized linear models (GLMs), the variance function describes:
Answer: How the variance of the response depends on its mean
The variance function V(μ) in a GLM specifies how the variance of the response relates to its mean, distinguishing different members of the exponential family.
Cook's distance for observation i is used to measure:
Answer: The overall influence of observation i on all fitted values
Cook's distance combines leverage and residual size to quantify how much the regression estimates change if observation i is removed.
The Neyman-Pearson lemma establishes that the most powerful test for a simple vs. simple hypothesis test is based on:
Answer: The likelihood ratio statistic
The Neyman-Pearson lemma proves that the likelihood ratio test (rejecting when L(H₁)/L(H₀) > k) is the uniformly most powerful test at level α.