Project Planning & Execution Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Project Planning & Execution flashcards as text
During the scoping phase of an ML project, a team discovers the labeled dataset has only 500 samples for a 10-class classification task. What is the MOST appropriate first response?
Answer: Re-evaluate feasibility and explore data augmentation or semi-supervised approaches
Re-evaluating feasibility and exploring augmentation or semi-supervised learning addresses the core data scarcity problem before committing to a training approach.
A stakeholder requests a real-time fraud detection system with sub-10ms latency. Which planning consideration is MOST critical to address early?
Answer: Defining the latency budget across all pipeline components including inference, pre-processing, and I/O
End-to-end latency budgeting must account for all pipeline stages, not just model inference, and must be validated early to avoid architectural rework.
Which artifact BEST serves as the single source of truth for tracking ML experiment reproducibility across a team?
Answer: An MLflow or similar experiment tracking registry logging code version, data hash, params, and metrics
Experiment tracking platforms like MLflow capture code version, data lineage, hyperparameters, and metrics in a queryable, reproducible format.
A project manager wants to estimate compute costs for training a large language model. Which factor has the GREATEST impact on total cost?
Answer: Model parameter count, dataset size, and number of training steps combined
Compute cost scales with the product of model size, data volume, and training steps, making these the dominant cost drivers.
When defining the success criteria for an ML project, what is the PRIMARY risk of using accuracy as the sole metric for an imbalanced dataset?
Answer: A model predicting only the majority class can achieve high accuracy while failing on the minority class
On imbalanced datasets, a trivial majority-class predictor can achieve misleadingly high accuracy, masking complete failure on the minority class.
In ML project execution, what does 'data versioning' primarily help prevent?
Answer: Silent dataset drift causing irreproducible experiments when the underlying data changes
Data versioning creates immutable snapshots so that experiments reference the exact dataset used, preventing silent reproducibility failures when source data is updated.
A team is planning sprint tasks for an ML project. Which task type should be placed in the EARLIEST sprint to reduce project risk?
Answer: Building a baseline end-to-end pipeline that touches all system components
An early end-to-end baseline pipeline exposes integration risks, latency issues, and data quality problems before significant investment is made in modeling.