โ† All AML Flashcard Decks

Technology & Digital Tools Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Technology & Digital Tools flashcards as text
  1. Which distributed computing framework is most commonly used for training large-scale ML models across multiple GPUs on a single node?

    Answer: PyTorch Distributed Data Parallel (DDP)

    PyTorch DDP synchronizes gradients across GPUs within a single node using an all-reduce operation, making it the standard for single-node multi-GPU training.

  2. In MLflow, what is the purpose of the 'Model Registry' component?

    Answer: To manage model versioning, staging, and production transitions

    MLflow Model Registry provides a centralized hub for managing the full lifecycle of ML models including versioning and stage transitions (Staging, Production, Archived).

  3. What is the primary advantage of using ONNX (Open Neural Network Exchange) format for ML models?

    Answer: It enables interoperability between different ML frameworks

    ONNX defines a common format so models trained in one framework (e.g., PyTorch) can be deployed using a different runtime or framework (e.g., TensorFlow, ONNX Runtime).

  4. When using Kubernetes for ML workload orchestration, what resource type is typically used to run a one-time batch training job?

    Answer: Job

    A Kubernetes Job creates one or more pods to run a task to completion and terminates them afterward, making it ideal for finite ML training runs.

  5. Which technique does TensorFlow's tf.data API primarily optimize?

    Answer: Data ingestion and preprocessing pipeline throughput

    tf.data optimizes the ETL pipeline feeding data into training by enabling parallelism, prefetching, and caching to prevent GPU starvation.

  6. In the context of feature stores (e.g., Feast, Tecton), what problem does the 'training-serving skew' refer to?

    Answer: Discrepancies between features computed offline for training and online for inference

    Training-serving skew occurs when the feature transformations applied at training time differ from those applied at inference time, causing degraded production performance.

  7. What does the '--mixed-precision' training flag (e.g., in Hugging Face Trainer) primarily achieve?

    Answer: It uses FP16 or BF16 for forward/backward passes to reduce memory usage and increase throughput

    Mixed precision training stores activations and gradients in lower-precision formats (FP16/BF16) while keeping master weights in FP32, cutting memory usage and speeding up computation on modern GPUs.

Technology & Digital Tools Flashcards โ€” AML Study Cards with Answers