← All MS-DS Master of Data science Flashcard Decks

Statistical and Probabilistic Analysis Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Statistical and Probabilistic Analysis flashcards as text
  1. Which nonparametric test is the rank-based alternative to a paired t-test?

    Answer: Wilcoxon signed-rank test

    The Wilcoxon signed-rank test analyzes paired differences using ranks, making it the nonparametric equivalent of the paired t-test.

  2. If X and Y are jointly distributed random variables, Cov(X, Y) = 0 implies independence only when:

    Answer: X and Y are jointly normally distributed

    Zero covariance implies independence only for jointly normal distributions; in general, uncorrelated does not mean independent.

  3. A 90% credible interval in Bayesian statistics differs from a 90% confidence interval in that:

    Answer: The credible interval directly states there is a 90% posterior probability the parameter lies in the interval

    A Bayesian credible interval has a direct probability interpretation: given the prior and data, P(θ ∈ interval) = 90%.

  4. In hypothesis testing, the p-value is defined as:

    Answer: The probability of obtaining a test statistic as extreme or more extreme than observed, assuming H₀ is true

    The p-value is a conditional probability computed under H₀; small p-values suggest the observed data are unlikely under the null.

  5. Which distribution is appropriate for modeling the time between events in a Poisson process?

    Answer: Exponential distribution

    In a Poisson process, inter-arrival times are exponentially distributed with rate λ equal to the Poisson arrival rate.

  6. A confusion matrix shows: TP=80, FP=20, FN=10, TN=90. What is the precision of the classifier?

    Answer: 0.80

    Precision = TP / (TP + FP) = 80 / (80 + 20) = 80/100 = 0.80.

  7. The F-statistic in ANOVA is the ratio of:

    Answer: Between-group variance to within-group variance

    F = Mean Square Between / Mean Square Within; a large F suggests group means differ more than expected by chance.