Technology & Digital Tools Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Technology & Digital Tools flashcards as text
Which Kubernetes resource is most appropriate for deploying a stateful ML model serving service that requires stable network identifiers?
Answer: StatefulSet
StatefulSets provide stable pod names, persistent storage, and ordered deployment, which is useful for services requiring consistent identity across restarts.
What is 'LoRA' (Low-Rank Adaptation) used for in large language model fine-tuning?
Answer: Fine-tuning only low-rank decomposition matrices instead of all model weights
LoRA freezes pretrained weights and injects trainable low-rank matrices into attention layers, drastically reducing the number of trainable parameters and GPU memory needed for fine-tuning.
In TensorFlow Serving, which protocol is preferred for high-throughput, low-latency model inference in production?
Answer: gRPC
gRPC uses Protocol Buffers and HTTP/2 for binary serialization and multiplexing, providing lower latency and higher throughput than REST for TF Serving inference requests.
What is the primary function of 'gradient checkpointing' in deep learning?
Answer: Trading compute for memory by recomputing activations during backpropagation instead of storing them
Gradient checkpointing discards intermediate activations during the forward pass and recomputes them during backpropagation, reducing peak memory usage at the cost of ~33% more computation.
Which format does the Hugging Face Hub use for efficient, sharded storage of large model weights that supports lazy loading?
Answer: Safetensors
Safetensors is a safe, fast format for storing tensors that supports memory-mapped lazy loading, avoids arbitrary code execution risks of pickle, and enables efficient sharding.
In the context of LLM inference optimization, what does 'KV cache' refer to?
Answer: Cached key-value attention matrices reused across autoregressive generation steps
The KV cache stores computed key and value tensors from previous tokens during autoregressive decoding, avoiding redundant recomputation and dramatically accelerating text generation.
What is 'online feature computation' in a feature store architecture?
Answer: Serving precomputed features from a low-latency store during real-time inference
Online feature computation serves previously computed and stored features from a low-latency store (e.g., Redis) at inference time, enabling sub-millisecond feature retrieval for real-time ML.