Evaluating Statistical Claims: Observational Studies and Experiments Flashcards
6 cards from real Bluebook SAT Test practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Evaluating Statistical Claims: Observational Studies and Experiments flashcards as text
A nutrition researcher wants to study whether eating breakfast daily improves academic performance in high school students. She surveys 500 students about their breakfast habits and then compares GPA data. Which of the following is the most significant limitation of this study design that prevents establishing causation?
Answer: Students who eat breakfast regularly may differ from those who don't in ways that also affect academic performance, such as household income or parental involvement.
This is an observational study, and the primary limitation is confounding variables. Students who eat breakfast regularly may come from more stable home environments, have higher family incomes, or have parents who are more involved in their education — all of which independently affect GPA. This confounding makes it impossible to isolate breakfast as the cause of any GPA difference. Sample size (A) is not the key issue here. GPA (C) is a widely accepted academic metric. Survey format (D) is irrelevant to the causation problem.
In a double-blind randomized controlled experiment testing a new anxiety medication, 200 participants are randomly assigned to either the medication or a placebo. After 12 weeks, the medication group shows a statistically significant improvement (p = 0.03). A researcher then notices that 30 participants in the medication group dropped out due to side effects, compared to only 5 in the placebo group. What is the most serious threat to the validity of the study's conclusion?
Answer: Differential dropout rates may mean the remaining medication group disproportionately excludes those who respond poorly, inflating the apparent benefit.
This is an example of attrition bias (also called differential dropout). When 30 participants left the medication group due to side effects versus only 5 in the placebo group, the remaining medication participants are not representative of those who started — they're likely the ones who tolerated the drug best. This systematically skews the results to look more favorable. The p-value of 0.03 (A) does meet the conventional threshold of 0.05. Study duration (B) is a separate concern not raised by the data. Placebo dosage (D) is irrelevant to inactive substances.
A study finds a strong positive correlation (r = 0.82) between the number of firefighters sent to a fire and the amount of property damage caused. A city councilmember argues this proves firefighters cause more damage and proposes reducing the fire department budget. Which statistical principle most directly explains the flaw in the councilmember's reasoning?
Answer: Correlation does not imply causation, and a lurking variable — fire severity — drives both the number of firefighters dispatched and the amount of damage.
This is a classic lurking variable scenario. Fire severity is the confounding factor that causes both more firefighters to be dispatched AND more property damage. Larger, more dangerous fires require more responders and naturally cause more destruction — there is no causal link from firefighters to damage. While (C) is technically true that an experiment would be needed for causation, the more precise and direct explanation is the lurking variable. The correlation of 0.82 (A) is indeed strong. Measurement scale (D) does not address the logical flaw.
Researchers conduct a randomized experiment in which participants are assigned to either a 'standing desk' group or a 'seated desk' group for 6 months. They measure back pain levels at the end. The study is described as 'single-blind.' Which of the following correctly identifies what 'single-blind' means in this context AND a resulting limitation?
Answer: Only the participants know their group assignment; this prevents demand characteristics but researchers may unconsciously bias their measurements.
In a study comparing standing vs. seated desks, it is physically impossible to blind participants — they obviously know whether they are standing or sitting. Therefore, 'single-blind' in this context most logically means participants are aware of their condition but the researchers measuring outcomes are not. However, this introduces a limitation: researchers who know which participants used which desk type may unconsciously rate or record pain levels in a biased direction (observer bias). Option A reverses the roles. Option C describes double-blind, which is impossible here. Option D describes analyst blinding only, which is a different design choice.
A social media company runs an A/B test to measure whether a new notification feature increases daily app usage. Users are randomly assigned to see the new feature (Group A) or the old interface (Group B). The test runs for 3 days, and Group A shows 14% more daily usage. The company concludes the feature is a success. Which of the following represents the most sophisticated critique of this conclusion?
Answer: Three days is insufficient to rule out novelty effect — users in Group A may show elevated engagement simply because the feature is new, not because it provides lasting value.
The novelty effect (also called the Hawthorne effect in some contexts) is the primary threat here. When users encounter a new feature, they often engage more in the short term simply out of curiosity. A 3-day window cannot distinguish between genuine long-term behavioral change and a temporary spike driven by novelty. Longer follow-up is needed to confirm the effect persists. Option A is not necessary since Group B serves as the control. 14% (C) would actually be quite meaningful at scale. Random assignment (D) is entirely achievable digitally — this is how all major platforms run A/B tests.
A poll of 1,200 randomly selected registered voters finds that 54% support a ballot measure, with a margin of error of ±3 percentage points at a 95% confidence level. A political analyst states: 'We can be 95% confident that exactly 54% of all registered voters support this measure.' What is wrong with the analyst's interpretation?
Answer: The confidence interval applies to a range (51%–57%), not a single value, and 95% confidence describes the reliability of the method across repeated samples, not certainty about this specific interval.
The analyst commits two related errors. First, the confidence interval is a range — 51% to 57% — not a single point estimate of 54%. Second, and more subtly, '95% confidence' does not mean 'we are 95% sure this specific interval contains the true value.' It means that if this polling method were repeated many times, 95% of the resulting intervals would contain the true population parameter. We cannot assign a probability to whether any one specific interval is correct — it either contains the true value or it doesn't. Sample size (A) of 1,200 is actually appropriate for a ±3% margin. Option B misdefines confidence level. Option D is incorrect — the margin applies symmetrically.