Statistical Inference Concepts Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Statistical Inference Concepts flashcards as text
In a chi-square goodness-of-fit test, the test statistic is large when:
Answer: Observed frequencies deviate substantially from expected frequencies
The chi-square statistic Σ(O−E)²/E is large when observed counts differ greatly from what the null hypothesis predicts.
The delta method is used to approximate the variance of g(X̄) when:
Answer: g is a smooth (differentiable) function and n is large
The delta method uses a first-order Taylor expansion of g around μ; it requires g to be differentiable and works well for large n via the CLT.
A p-value of 0.001 in a study with n = 10,000 primarily reflects:
Answer: A very small effect size that is detectable only due to the huge sample
With n = 10,000, even tiny effects produce very small p-values; the p-value conflates effect size with sample size.
In hypothesis testing, which error is controlled directly by setting the significance level α?
Answer: Type I error (rejecting a true H₀)
The significance level α is defined as the maximum acceptable probability of a Type I error (false positive).
The Rao-Blackwell theorem states that if T is a sufficient statistic and U is an unbiased estimator, then E[U|T] is:
Answer: Unbiased and has variance no greater than that of U
Conditioning an unbiased estimator on a sufficient statistic yields an unbiased estimator with equal or lower variance — the Rao-Blackwell improvement.
Which of the following correctly describes the relationship between a two-sided 95% confidence interval and a two-sided hypothesis test at α = 0.05?
Answer: H₀: μ = μ₀ is rejected at α = 0.05 if and only if μ₀ falls outside the 95% CI
There is a duality: a value μ₀ is in the 95% CI if and only if the test of H₀: μ = μ₀ fails to reject at α = 0.05.
A Bayesian credible interval differs from a frequentist confidence interval in that:
Answer: A 95% credible interval has a 95% posterior probability of containing the true parameter
A Bayesian credible interval directly states that P(θ ∈ I | data) = 0.95, a direct probability statement about the parameter given the data.