โ† All AML Flashcard Decks

Technology & Digital Tools Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Technology & Digital Tools flashcards as text
  1. Which Kubernetes resource is most appropriate for deploying a stateful ML model serving service that requires stable network identifiers?

    Answer: StatefulSet

    StatefulSets provide stable pod names, persistent storage, and ordered deployment, which is useful for services requiring consistent identity across restarts.

  2. What is 'LoRA' (Low-Rank Adaptation) used for in large language model fine-tuning?

    Answer: Fine-tuning only low-rank decomposition matrices instead of all model weights

    LoRA freezes pretrained weights and injects trainable low-rank matrices into attention layers, drastically reducing the number of trainable parameters and GPU memory needed for fine-tuning.

  3. In TensorFlow Serving, which protocol is preferred for high-throughput, low-latency model inference in production?

    Answer: gRPC

    gRPC uses Protocol Buffers and HTTP/2 for binary serialization and multiplexing, providing lower latency and higher throughput than REST for TF Serving inference requests.

  4. What is the primary function of 'gradient checkpointing' in deep learning?

    Answer: Trading compute for memory by recomputing activations during backpropagation instead of storing them

    Gradient checkpointing discards intermediate activations during the forward pass and recomputes them during backpropagation, reducing peak memory usage at the cost of ~33% more computation.

  5. Which format does the Hugging Face Hub use for efficient, sharded storage of large model weights that supports lazy loading?

    Answer: Safetensors

    Safetensors is a safe, fast format for storing tensors that supports memory-mapped lazy loading, avoids arbitrary code execution risks of pickle, and enables efficient sharding.

  6. In the context of LLM inference optimization, what does 'KV cache' refer to?

    Answer: Cached key-value attention matrices reused across autoregressive generation steps

    The KV cache stores computed key and value tensors from previous tokens during autoregressive decoding, avoiding redundant recomputation and dramatically accelerating text generation.

  7. What is 'online feature computation' in a feature store architecture?

    Answer: Serving precomputed features from a low-latency store during real-time inference

    Online feature computation serves previously computed and stored features from a low-latency store (e.g., Redis) at inference time, enabling sub-millisecond feature retrieval for real-time ML.