โ† All Data Science Flashcard Decks

Data Science Statistical Concepts and Analysis Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Science Statistical Concepts and Analysis Questions and Answers flashcards as text
  1. What does the coefficient of determination (R-squared) value of 0.85 indicate in a regression model?

    Answer: 85% of the variance in the dependent variable is explained by the model

    An R-squared of 0.85 means the independent variables in the model explain 85% of the variance in the dependent variable.

  2. A data scientist applies the Bonferroni correction when conducting 20 simultaneous hypothesis tests at alpha = 0.05. What is the adjusted significance level for each individual test?

    Answer: 0.0025

    The Bonferroni correction divides the overall significance level by the number of tests: 0.05 / 20 = 0.0025.

  3. Which of the following best describes a Type II error in statistical testing?

    Answer: Failing to reject a false null hypothesis

    A Type II error occurs when we fail to reject the null hypothesis even though it is actually false, meaning we miss a real effect.

  4. When applying Principal Component Analysis (PCA), what does the first principal component represent?

    Answer: The direction of maximum variance in the data

    The first principal component captures the direction (linear combination of features) along which the data exhibits the greatest variance.

  5. A bootstrap sampling procedure draws 1,000 samples of size n with replacement from the original dataset. What is the primary purpose of this technique?

    Answer: To estimate the sampling distribution of a statistic

    Bootstrapping repeatedly resamples the data to approximate the sampling distribution of a statistic, enabling confidence interval estimation without parametric assumptions.

  6. In Bayesian statistics, what does the posterior distribution represent?

    Answer: The updated belief about parameters after observing data

    The posterior distribution combines the prior belief with the observed data (via the likelihood) to produce an updated probability distribution over the parameters.