Quality Assurance & Improvement Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Quality Assurance & Improvement flashcards as text
Which cross-validation strategy is most appropriate when your dataset has significant temporal ordering?
Answer: Time-series split (walk-forward)
Time-series split (walk-forward validation) prevents data leakage by always training on past data and validating on future data.
A model achieves 99% accuracy on an imbalanced dataset where 99% of samples are class 0. What is the most important additional metric to evaluate?
Answer: Matthews Correlation Coefficient (MCC)
MCC accounts for all four confusion matrix values and provides a reliable metric even when classes are highly imbalanced.
What does a calibration curve (reliability diagram) measure in a classifier?
Answer: Alignment between predicted probabilities and actual outcome frequencies
A calibration curve plots predicted probabilities against actual frequencies to show whether a model's confidence scores are trustworthy.
You observe that validation loss stops improving after epoch 20 but training loss keeps decreasing. The correct QA response is to:
Answer: Apply early stopping at epoch 20 and regularize the model
The divergence between training and validation loss is a clear overfitting signal; early stopping and regularization directly address it.
Which technique quantifies uncertainty by training multiple models on bootstrap samples of the training data?
Answer: Bootstrap aggregating (Bagging)
Bagging trains multiple models on bootstrapped datasets; the variance of their predictions estimates model uncertainty.
During error analysis, you find that 80% of misclassifications occur on one specific data slice. The best QA action is to:
Answer: Upsample that slice and retrain with targeted augmentation
Targeted upsampling and augmentation of the underperforming slice directly addresses the root cause of slice-specific failures.
The Brier Score is used to evaluate:
Answer: Calibration and sharpness of probabilistic forecasts
The Brier Score is the mean squared error of probability predictions, measuring both calibration and sharpness of probabilistic classifiers.