Data Science Model Performance and Evaluation Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Model Performance and Evaluation Questions and Answers flashcards as text
Which metric is most appropriate for evaluating a classification model when the dataset has a severe class imbalance?
Answer: Precision-Recall AUC
Precision-Recall AUC focuses on the minority class performance and is far more informative than accuracy when classes are heavily imbalanced.
What does a high variance and low bias in a model typically indicate?
Answer: The model is overfitting the training data
High variance with low bias means the model fits training data very closely but fails to generalize, which is the hallmark of overfitting.
In k-fold cross-validation with k=10, how many times is each data point used in the test set?
Answer: 1 time
In k-fold cross-validation, the data is split into k folds, and each fold serves as the test set exactly once across all iterations.
What is the primary purpose of a calibration curve (reliability diagram) in model evaluation?
Answer: To assess whether predicted probabilities match actual outcome frequencies
A calibration curve plots predicted probabilities against actual frequencies to reveal whether a model's confidence scores are trustworthy.
Which evaluation approach is most suitable when you need to assess a time-series forecasting model?
Answer: Walk-forward validation
Walk-forward validation respects temporal ordering by always training on past data and testing on future data, preventing data leakage.
What does a Matthews Correlation Coefficient (MCC) of 0 indicate about a binary classifier?
Answer: Performance no better than random guessing
An MCC of 0 indicates the classifier performs no better than random prediction, while +1 indicates perfect and -1 indicates total disagreement.