← All MS-DS Master of Data science Flashcard Decks

Statistical Inference & Regression Models Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Statistical Inference & Regression Models flashcards as text
  1. In multiple regression, the partial F-test is used to:

    Answer: Compare a reduced model to a full model by testing a subset of predictors simultaneously

    The partial F-test evaluates whether a group of predictors collectively adds significant explanatory power beyond what the reduced model already explains.

  2. A Q-Q plot of regression residuals that shows heavy tails indicates:

    Answer: Departure from normality with more extreme values than a normal distribution

    Heavy tails in a Q-Q plot mean the residuals have a leptokurtic distribution with more extreme outliers than expected under normality.

  3. Which scenario correctly describes an interaction term in multiple regression?

    Answer: The effect of X₁ on Y depends on the value of X₂

    An interaction term (X₁ × X₂) captures the situation where the marginal effect of one predictor on the response changes depending on another predictor's value.

  4. The bootstrap method is used in regression to:

    Answer: Estimate the sampling distribution of an estimator by resampling with replacement

    Bootstrap resampling creates many pseudo-samples from the observed data to empirically estimate the sampling distribution and standard errors of estimators.

  5. In generalized linear models (GLMs), the variance function describes:

    Answer: How the variance of the response depends on its mean

    The variance function V(μ) in a GLM specifies how the variance of the response relates to its mean, distinguishing different members of the exponential family.

  6. Cook's distance for observation i is used to measure:

    Answer: The overall influence of observation i on all fitted values

    Cook's distance combines leverage and residual size to quantify how much the regression estimates change if observation i is removed.

  7. The Neyman-Pearson lemma establishes that the most powerful test for a simple vs. simple hypothesis test is based on:

    Answer: The likelihood ratio statistic

    The Neyman-Pearson lemma proves that the likelihood ratio test (rejecting when L(H₁)/L(H₀) > k) is the uniformly most powerful test at level α.