Technology & Digital Tools Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Technology & Digital Tools flashcards as text
In Apache Spark MLlib, which abstraction replaced the original RDD-based API as the primary interface for machine learning?
Answer: DataFrame-based ML Pipelines
Spark MLlib's DataFrame-based Pipeline API (spark.ml) replaced the older RDD-based API (spark.mllib) for better integration with Spark SQL and improved performance.
Which tool is specifically designed for hyperparameter optimization and supports population-based training, Bayesian optimization, and early stopping?
Answer: Ray Tune
Ray Tune is a scalable hyperparameter tuning library that integrates multiple search algorithms (Bayesian, PBT, ASHA) and runs trials in parallel across a cluster.
What is 'model quantization' in the context of ML model optimization?
Answer: Converting model weights from higher-precision to lower-precision data types
Quantization reduces model size and inference latency by representing weights and activations in lower-bit formats (e.g., INT8 instead of FP32) with minimal accuracy loss.
In a typical CI/CD pipeline for ML (MLOps), what step immediately follows model training and precedes deployment?
Answer: Model evaluation and validation
After training, the model must pass evaluation gates (accuracy thresholds, fairness checks, performance benchmarks) before it is approved for deployment.
Which GPU memory optimization technique allows training models larger than a single GPU's VRAM by partitioning model layers across multiple GPUs?
Answer: Pipeline Parallelism (Model Parallelism)
Pipeline/model parallelism splits the model's layers across multiple GPUs so each GPU holds and computes only a portion of the network.
What is the role of a 'vector database' (e.g., Pinecone, Weaviate, Chroma) in modern ML systems?
Answer: Efficient similarity search over high-dimensional embedding vectors
Vector databases index high-dimensional embeddings and support approximate nearest-neighbor (ANN) search, enabling fast semantic retrieval for RAG and recommendation systems.
In Weights & Biases (W&B), what does 'wandb.watch()' do during model training?
Answer: Logs model gradients and parameter histograms to the W&B dashboard
wandb.watch() hooks into the model to automatically log gradients and parameter distributions each step, helping detect issues like vanishing/exploding gradients.