Quality Assurance & Improvement Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Quality Assurance & Improvement flashcards as text
Which statistical test is most appropriate for detecting covariate shift between training and production data distributions?
Answer: Two-sample Kolmogorov-Smirnov test
The two-sample KS test compares the empirical CDFs of two continuous distributions without assuming normality, making it ideal for detecting feature drift.
Concept drift occurs when:
Answer: The relationship P(Y|X) between inputs and outputs changes over time
Concept drift specifically refers to a change in the conditional distribution P(Y|X), meaning the same inputs should now map to different outputs.
A production model's average prediction score drops significantly with no change in input feature distributions. This most likely indicates:
Answer: A data pipeline serialization error
When input distributions are stable but prediction scores change, the issue is typically in data pipeline processing or feature engineering, not the model itself.
Which monitoring metric is specifically designed to detect when a model's predictions become less correlated with eventual ground truth labels before those labels are available?
Answer: Population Stability Index (PSI)
PSI measures how much the distribution of model scores has shifted from a reference period, acting as an early proxy for performance degradation.
In shadow mode deployment for ML QA, the shadow model's predictions are:
Answer: Logged and evaluated but never surfaced to end users
Shadow mode runs the candidate model in parallel, capturing predictions for offline evaluation without any user-facing impact.
When setting up automated model retraining triggers in a production ML system, which signal is most directly actionable?
Answer: Statistically significant drop in held-out evaluation metric
A statistically significant performance degradation on a labeled evaluation set directly indicates the model needs retraining with updated data.
What is the primary risk of using a static holdout test set for ongoing model quality monitoring over multiple retraining cycles?
Answer: Test set contamination through repeated model selection on the same data
Repeated selection of models based on the same test set causes the test set to implicitly influence model development, leaking information and inflating performance estimates.