โ† All MS-DS Master of Data science Flashcard Decks

Research & Data Analysis Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Research & Data Analysis flashcards as text
  1. A researcher wants to determine whether a new teaching method improves test scores compared to a traditional method. Which study design is most appropriate?

    Answer: Randomized controlled experiment

    A randomized controlled experiment allows causal inference by randomly assigning participants to treatment and control groups.

  2. In hypothesis testing, a p-value of 0.03 with a significance level of 0.05 means:

    Answer: Reject the null hypothesis

    Since 0.03 < 0.05, the result is statistically significant and we reject the null hypothesis.

  3. Which technique is used to assess the stability of a regression model by partitioning data into training and validation subsets multiple times?

    Answer: K-fold cross-validation

    K-fold cross-validation splits data into k subsets and iteratively trains/validates across all folds to estimate model performance.

  4. A dataset has a mean of 50, median of 45, and mode of 40. This distribution is best described as:

    Answer: Positively skewed

    When mean > median > mode, the distribution has a longer right tail, indicating positive (right) skew.

  5. A data scientist uses Lasso regression instead of OLS. The primary advantage of Lasso in high-dimensional data is:

    Answer: It performs automatic feature selection by shrinking some coefficients to zero

    Lasso's L1 penalty shrinks some coefficients to exactly zero, effectively selecting a sparse subset of predictors.

  6. In a meta-analysis, publication bias most commonly results in:

    Answer: Overestimation of effect sizes because significant results are more likely published

    Studies with statistically significant results are more likely published, inflating the pooled effect size in meta-analyses.

  7. Which measure of association is most appropriate when analyzing the relationship between two continuous variables that may not have a linear relationship?

    Answer: Spearman rank correlation

    Spearman rank correlation captures monotonic relationships and is robust to non-linearity and outliers unlike Pearson correlation.