Model Evaluation & Optimization Techniques Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Model Evaluation & Optimization Techniques flashcards as text
What is the key insight behind cyclical learning rate (CLR) schedules compared to monotonically decaying schedules?
Answer: Periodically increasing the learning rate can escape saddle points and sharp minima, potentially finding flatter optima
CLR's periodic increases temporarily escape sharp, narrow minima associated with poor generalization, helping find flatter loss landscape regions that generalize better.
When evaluating a regression model, which metric is scale-independent and allows comparison across datasets with different target variable magnitudes?
Answer: Mean Absolute Percentage Error (MAPE)
MAPE expresses errors as a percentage of actual values, making it scale-independent and comparable across datasets with different target magnitudes, unlike absolute metrics.
In the context of Neural Architecture Search (NAS), what does the 'weight sharing' strategy in one-shot NAS achieve?
Answer: It trains a single supernet where all candidate architectures share parameters, dramatically reducing search cost
One-shot NAS trains a supernet once where all sub-architectures share parameters, then evaluates candidates by sampling from the supernet, reducing search cost from thousands of GPU-days to hours.
Which statistical test is recommended for comparing multiple ML models across multiple datasets to avoid inflated Type I error?
Answer: Friedman test followed by post-hoc Nemenyi test
The Friedman test is a non-parametric rank-based test for multiple classifiers over multiple datasets, and the Nemenyi post-hoc test identifies which pairs differ significantly.
What is the purpose of temperature scaling in post-hoc model calibration?
Answer: It adjusts the softmax temperature using a single learned scalar to align predicted probabilities with empirical frequencies
Temperature scaling divides logits by a learned temperature T before softmax, uniformly adjusting confidence levels to better match empirical accuracy without changing predictions.
In population-based training (PBT), what mechanism allows hyperparameters to evolve during training rather than being fixed at initialization?
Answer: Periodic exploitation (copying weights from better-performing workers) and exploration (perturbing hyperparameters)
PBT runs a population of workers in parallel, periodically replacing underperforming workers' weights with top performers' weights while randomly perturbing their hyperparameters.
Which evaluation protocol is most appropriate for measuring a recommender system's performance when the goal is relevance of the top-N items shown to users?
Answer: Normalized Discounted Cumulative Gain (NDCG@N)
NDCG@N measures the quality of the top-N ranked items, discounting relevance by position, making it directly aligned with the user experience of seeing a ranked recommendation list.