โ† All Data Science Flashcard Decks

Model Performance and Evaluation Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Model Performance and Evaluation flashcards as text
  1. Stratified k-fold cross-validation differs from standard k-fold by:

    Answer: Preserving the class distribution in each fold

    Stratified folds maintain the original class proportions in every fold, important for imbalanced data.

  2. Specificity measures a classifier's ability to correctly identify:

    Answer: Actual negatives

    Specificity (true negative rate) is the proportion of actual negatives correctly identified.

  3. If two models have identical accuracy but different ROC-AUC, the higher-AUC model is better at:

    Answer: Ranking positives above negatives across thresholds

    AUC reflects ranking quality across all thresholds, independent of a single decision cutoff.

  4. A learning curve where both training and validation error remain high and close together indicates:

    Answer: Underfitting (high bias)

    High, converged errors signal underfitting, where the model is too simple for the data.

  5. Cohen's Kappa is preferred over raw accuracy because it:

    Answer: Accounts for agreement expected by chance

    Cohen's Kappa corrects observed accuracy for the agreement expected by random chance.

  6. When reporting model performance, why is a confidence interval on the metric valuable?

    Answer: It quantifies uncertainty in the estimate due to finite test data

    Confidence intervals communicate how much the metric might vary given limited test data.

  7. Nested cross-validation is used primarily to:

    Answer: Obtain an unbiased performance estimate while tuning hyperparameters

    Nested CV separates hyperparameter tuning (inner loop) from performance estimation (outer loop) to avoid optimistic bias.