Model Performance and Evaluation Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Model Performance and Evaluation flashcards as text
A spam classifier flags 95 of 100 actual spam emails but also flags 40 legitimate emails as spam. Which metric most directly reflects the cost of those false alarms?
Answer: Precision
Precision penalizes false positives, capturing the cost of legitimate emails wrongly flagged as spam.
On a dataset where 99% of cases are negative, a model that always predicts 'negative' achieves 99% accuracy. What does this illustrate?
Answer: The accuracy paradox on imbalanced data
The accuracy paradox shows accuracy can be misleadingly high on imbalanced data despite a useless model.
Which metric is the harmonic mean of precision and recall?
Answer: F1 score
The F1 score is the harmonic mean of precision and recall, balancing both.
A regression model has a low training error but a much higher test error. This is a sign of:
Answer: Overfitting
A large gap with low training error and high test error indicates overfitting (high variance).
The area under the ROC curve (AUC) represents the probability that the model ranks a randomly chosen positive higher than a randomly chosen negative. An AUC of 0.5 means:
Answer: No better than random guessing
An AUC of 0.5 corresponds to random ranking with no discriminative ability.
Which evaluation approach gives a more reliable performance estimate on small datasets by averaging across multiple train/test splits?
Answer: k-fold cross-validation
k-fold cross-validation averages performance over multiple folds, reducing variance of the estimate.
In a confusion matrix, recall (sensitivity) is calculated as:
Answer: TP / (TP + FN)
Recall is true positives divided by all actual positives, TP / (TP + FN).