Advanced Machine Learning (AML) Certification — Questions and Answers
Question 1: In data reporting, what is the primary advantage of using a waterfall chart over a simple bar chart?
- More effective for time-series trends
- Clearly shows cumulative effect of sequential positive and negative changes (Correct answer)
- Easier to read for non-technical audiences
- Better for showing proportions
Correct answer: Clearly shows cumulative effect of sequential positive and negative changes
Waterfall charts decompose a total into sequential incremental contributions, making it ideal for reporting how components add to or subtract from a KPI.
Question 2: What is the primary purpose of feature scaling in machine learning?
- To remove outliers from the dataset
- To encode categorical variables
- To ensure all features contribute equally to model training (Correct answer)
- To reduce the number of features
Correct answer: To ensure all features contribute equally to model training
Feature scaling normalizes feature ranges so that no single feature dominates model training due to its magnitude.
Question 3: What is shadow deployment in machine learning operations?
- Testing a model with synthetic data before live deployment
- Running a new model in parallel with the production model without using its predictions for actual decisions (Correct answer)
- Running a model only during off-peak hours to save resources
- Deploying a model to a private server without public access
Correct answer: Running a new model in parallel with the production model without using its predictions for actual decisions
Shadow deployment routes live traffic to both the current and new model, logging new model outputs for evaluation without affecting end users.
Question 4: Which technique handles missing values by replacing them with the average of the column?
- Label encoding
- One-hot encoding
- Standardization
- Mean imputation (Correct answer)
Correct answer: Mean imputation
Mean imputation replaces missing values with the column's average, preserving the dataset's overall distribution.
Question 5: Which technique generates synthetic minority class samples to address class imbalance?
- Bootstrapping
- Random undersampling
- SMOTE (Correct answer)
- Stratified sampling
Correct answer: SMOTE
SMOTE creates new synthetic samples by interpolating between existing minority class examples rather than duplicating them.
Question 6: What is model drift?
- The movement of model weights during training
- A regularization technique for production models
- The degradation of model performance as real-world data distribution changes from training data (Correct answer)
- The gradual improvement of a model over time
Correct answer: The degradation of model performance as real-world data distribution changes from training data
Model drift occurs when the statistical properties of input data or target relationships change post-deployment, reducing model accuracy.
Question 7: What is the purpose of a train-validation-test split in model development?
- To perform feature engineering
- To increase the size of the training set
- To tune hyperparameters and evaluate final model performance on unseen data (Correct answer)
- To balance class distributions
Correct answer: To tune hyperparameters and evaluate final model performance on unseen data
The validation set enables hyperparameter tuning while the test set provides an unbiased final performance estimate.
Question 8: Which method identifies predictive features by measuring how much each reduces impurity in a trained Random Forest?
- Correlation matrix analysis
- Recursive Feature Elimination
- Variance Inflation Factor (VIF)
- Feature importance from Random Forest (Correct answer)
Correct answer: Feature importance from Random Forest
Random Forest models compute feature importance scores based on average impurity reduction contributed by each feature across all trees.
Question 9: What is the difference between a regulation and a standard in AML practice?
- Standards are always stricter than regulations
- Regulations are legally binding requirements; standards are voluntary guidelines that may become regulatory through adoption (Correct answer)
- Regulations apply only to individuals, standards to organizations
- They are exactly the same thing
Correct answer: Regulations are legally binding requirements; standards are voluntary guidelines that may become regulatory through adoption
Regulations are legally enforceable rules issued by government agencies, while standards are developed by industry bodies and are typically voluntary. However, standards often become de facto requirements through regulatory adoption or contractual obligations.
Question 10: Which component adjusts the weights during training in neural networks?
- Learning rate
- Optimizer (Correct answer)
- Loss function
- Activation function
Correct answer: Optimizer
The optimizer is the component in neural networks responsible for adjusting the weights and biases during the training process. It uses the gradients calculated by backpropagation to determine how to update these parameters in a way that minimizes the loss function. Common optimizers include Stochastic Gradient Descent (SGD), Adam, and RMSprop, each with different strategies for navigating the loss landscape.
Question 11: A global company seeks to transfer EU personal data used for ML training to a US data center. The most legally robust mechanism under current GDPR rules is:
- Relying on the now-invalidated Privacy Shield agreement
- Encrypting data at rest without any additional legal mechanism
- Using Standard Contractual Clauses (SCCs) combined with a Transfer Impact Assessment (TIA) (Correct answer)
- Obtaining a one-time consent from each data subject
Correct answer: Using Standard Contractual Clauses (SCCs) combined with a Transfer Impact Assessment (TIA)
Following Schrems II, SCCs supplemented by a Transfer Impact Assessment (TIA) are the primary compliant mechanism for transferring personal data from the EU to the US for ML workloads.
Question 12: What does the Variance Inflation Factor (VIF) measure in regression analysis?
- The importance of features in a regression model
- The degree of multicollinearity among predictor variables (Correct answer)
- The inflation of model error as more features are added
- The variance of individual feature distributions
Correct answer: The degree of multicollinearity among predictor variables
VIF quantifies how much the variance of a regression coefficient is inflated due to correlation with other predictor variables.
Question 13: In a GAN, what condition describes the theoretical equilibrium where the generator perfectly replicates the data distribution?
- The generator loss reaches zero
- The discriminator correctly classifies 100% of real images
- The discriminator outputs 0.5 for all inputs (Correct answer)
- The gradient penalty term converges to one
Correct answer: The discriminator outputs 0.5 for all inputs
At Nash equilibrium the discriminator cannot distinguish real from generated samples, outputting 0.5 (random chance) for every input.
Question 14: Which tool is commonly used to containerize ML models for consistent deployment across environments?
- NumPy
- Docker (Correct answer)
- Scikit-learn
- Jupyter Notebook
Correct answer: Docker
Docker packages ML models with their dependencies into portable containers that run consistently across development and production environments.
Question 15: In multi-objective hyperparameter optimization, what does the Pareto frontier represent?
- The single best configuration across all objectives
- The average performance weighted by objective importance
- The set of configurations where no objective can be improved without degrading another (Correct answer)
- The region of hyperparameter space with highest variance
Correct answer: The set of configurations where no objective can be improved without degrading another
The Pareto frontier contains all non-dominated solutions — configurations where improving one objective (e.g., accuracy) necessarily worsens another (e.g., inference latency).
Question 16: Equalized odds as a fairness constraint requires that a model:
- Achieve equal true positive and false positive rates across groups (Correct answer)
- Produce the same positive prediction rate across all demographic groups
- Minimize the maximum error rate across any single group
- Use identical features for all demographic groups
Correct answer: Achieve equal true positive and false positive rates across groups
Equalized odds requires both true positive rate (sensitivity) and false positive rate to be equal across protected attribute groups.
Question 17: What does one-hot encoding do to categorical variables?
- Creates binary columns for each category (Correct answer)
- Scales categories to a 0-1 range
- Assigns numeric labels to categories
- Removes low-frequency categories
Correct answer: Creates binary columns for each category
One-hot encoding converts each category into a separate binary column, avoiding the assumption of ordinal relationships.
Question 18: What is the curse of dimensionality in machine learning?
- The problem of having too many training samples
- The difficulty of visualizing high-dimensional data
- The computational cost of training deep neural networks
- The phenomenon where data becomes sparse as dimensions increase, degrading model performance (Correct answer)
Correct answer: The phenomenon where data becomes sparse as dimensions increase, degrading model performance
As feature dimensions increase, data points become increasingly sparse, making distance-based algorithms less effective.
Question 19: Which NLP task involves assigning a label (positive, negative, neutral) to a piece of text based on its emotional tone?
- Sentiment Analysis (Correct answer)
- Machine Translation
- Named Entity Recognition
- Coreference Resolution
Correct answer: Sentiment Analysis
Sentiment analysis classifies text according to the opinion or emotion expressed, commonly as positive, negative, or neutral.
Question 20: Which US law most directly regulates the use of ML models in consumer credit reporting and adverse action notices?
- Fair Credit Reporting Act (FCRA) (Correct answer)
- Bank Secrecy Act (BSA)
- Dodd-Frank Act Section 1071
- Gramm-Leach-Bliley Act (GLBA)
Correct answer: Fair Credit Reporting Act (FCRA)
FCRA governs how consumer report information is collected, used, and shared, and requires adverse action notices with reasons when ML-driven credit decisions negatively affect consumers.
Question 21: What is the purpose of model versioning in MLOps?
- To label different datasets used to train a model
- To track distinct model iterations with metadata to enable rollback and comparison (Correct answer)
- To compress model weights for faster inference
- To increment model complexity with each training run
Correct answer: To track distinct model iterations with metadata to enable rollback and comparison
Model versioning logs each trained model artifact with its configuration, metrics, and data lineage to support governance and rollback.
Question 22: In NLP, what does Named Entity Recognition (NER) identify in text?
- Real-world entities such as people, organizations, and locations (Correct answer)
- The overall sentiment of a document
- Duplicate sentences across a corpus
- The grammatical role of each word in a sentence
Correct answer: Real-world entities such as people, organizations, and locations
NER tags spans of text that refer to specific entities like person names, companies, dates, and geographic locations.
Question 23: Concept drift occurs when:
- The model's weights decay due to long deployment
- The relationship P(Y|X) between inputs and outputs changes over time (Correct answer)
- Training batch size becomes inconsistent
- The input feature distribution P(X) changes over time
Correct answer: The relationship P(Y|X) between inputs and outputs changes over time
Concept drift specifically refers to a change in the conditional distribution P(Y|X), meaning the same inputs should now map to different outputs.
Question 24: Why is k-fold cross-validation used when evaluating feature engineering choices?
- To automatically select the optimal number of features
- To handle class imbalance
- To obtain a more reliable performance estimate across multiple data subsets (Correct answer)
- To increase the training dataset size
Correct answer: To obtain a more reliable performance estimate across multiple data subsets
K-fold cross-validation rotates the validation set across k subsets, giving a more robust estimate than a single train-test split.
Question 25: Which feature engineering technique creates new features from products or powers of existing features?
- Polynomial feature expansion (Correct answer)
- Feature scaling
- Label encoding
- Imputation
Correct answer: Polynomial feature expansion
Polynomial feature expansion generates new features as products or powers of existing ones, enabling linear models to capture non-linear patterns.
Question 26: What is the role of an ML model registry in the MLOps lifecycle?
- To automate hyperparameter optimization across training runs
- To store raw training datasets and feature CSVs
- To serve as a central catalog for tracking, versioning, and managing trained model artifacts through the deployment lifecycle (Correct answer)
- To monitor live model predictions and alert on drift
Correct answer: To serve as a central catalog for tracking, versioning, and managing trained model artifacts through the deployment lifecycle
A model registry provides a governed repository where teams can register, version, stage, and promote models from development to production.
Question 27: What is semantic segmentation in computer vision?
- Assigning a class label to every pixel in an image (Correct answer)
- Detecting and localizing objects with bounding boxes
- Generating a text caption for an image
- Ranking image regions by confidence score
Correct answer: Assigning a class label to every pixel in an image
Semantic segmentation classifies each pixel of an image into a category, producing a pixel-wise class map of the entire scene.
Question 28: A company wants to deploy an ML model under a strict regulatory compliance requirement. Which planning step is UNIQUELY critical compared to a non-regulated deployment?
- Switching from Python to R for statistical compliance
- Increasing the size of the validation set
- Using a GPU cluster instead of CPU training
- Documenting model lineage, maintaining audit trails, and planning for explainability requirements before development begins (Correct answer)
Correct answer: Documenting model lineage, maintaining audit trails, and planning for explainability requirements before development begins
Regulated environments require audit trails, explainability artifacts, and lineage documentation that must be architected into the system from the start, not added retroactively.
Question 29: In AdaBoost, what happens to the weights of incorrectly classified samples after each boosting round?
- Their weights are decreased so the next classifier focuses on different samples
- Their weights remain unchanged; only the classifier weights change
- Their weights are reset to uniform to prevent overfitting
- Their weights are increased so the next classifier prioritizes them (Correct answer)
Correct answer: Their weights are increased so the next classifier prioritizes them
AdaBoost increases the sample weights of misclassified instances, forcing subsequent weak learners to focus more attention on the harder-to-classify examples.
Question 30: What is the BLEU score used to measure in NLP?
- Classification accuracy on text labels
- Number of out-of-vocabulary tokens in a corpus
- Semantic similarity between two sentence embeddings
- Quality of machine-generated text by comparing n-gram overlap with reference translations (Correct answer)
Correct answer: Quality of machine-generated text by comparing n-gram overlap with reference translations
BLEU measures how many n-grams in a machine translation match those in one or more reference translations, normalized by length.
Question 31: What is a feature store in an MLOps architecture?
- A cloud storage bucket for raw datasets
- A tool for automated feature selection
- A replica of the training database used for inference
- A centralized repository for storing, sharing, and serving precomputed features for ML models (Correct answer)
Correct answer: A centralized repository for storing, sharing, and serving precomputed features for ML models
Feature stores provide consistent, reusable features across training and serving pipelines, eliminating training-serving skew.
Question 32: Which outlier detection method uses the interquartile range (IQR) to flag extreme values?
- Z-score normalization
- Tukey's fence method (Correct answer)
- Min-max scaling
- Winsorization
Correct answer: Tukey's fence method
Tukey's fence method flags values beyond 1.5x IQR from Q1 or Q3 as outliers for removal or treatment.
Question 33: What is the primary benefit of feature selection in a machine learning pipeline?
- Reduces overfitting and training time (Correct answer)
- Improves data imputation accuracy
- Increases model complexity
- Adds synthetic training data
Correct answer: Reduces overfitting and training time
Feature selection removes irrelevant or redundant features, reducing overfitting risk and computational cost.
Question 34: What is the purpose of subword tokenization methods like Byte-Pair Encoding (BPE) used in models like GPT?
- Speed up inference by reducing vocabulary size to single characters
- Convert text to binary representations for efficient storage
- Replace all punctuation tokens before model training
- Balance vocabulary size and out-of-vocabulary handling by splitting rare words into frequent subword units (Correct answer)
Correct answer: Balance vocabulary size and out-of-vocabulary handling by splitting rare words into frequent subword units
BPE iteratively merges frequent character pairs into subword units, allowing models to handle rare and unseen words by decomposing them.
Question 35: Which statement correctly describes the bias-variance tradeoff in the context of k-Nearest Neighbors (kNN)?
- Larger k increases variance and decreases bias
- The choice of k does not affect bias or variance
- Smaller k decreases variance and increases bias
- Larger k increases bias and decreases variance (Correct answer)
Correct answer: Larger k increases bias and decreases variance
A larger k smooths the decision boundary by averaging more neighbors, increasing bias (underfitting) but reducing variance (sensitivity to noise).
Question 36: Which regularization technique explicitly prevents co-adaptation of neurons by randomly zeroing activations during training?
- Batch normalization
- Dropout (Correct answer)
- Label smoothing
- L2 weight decay
Correct answer: Dropout
Dropout randomly deactivates neurons during each training step, forcing the network to learn redundant representations and preventing co-adaptation.
Question 37: Which automated retraining trigger strategy is considered a best practice in production MLOps?
- Scheduled or event-driven retraining triggered by performance degradation thresholds or new data availability (Correct answer)
- Manual retraining only on a fixed monthly schedule
- Retraining only when a new model architecture is designed
- Retraining after every new user request to the API
Correct answer: Scheduled or event-driven retraining triggered by performance degradation thresholds or new data availability
Automated retraining pipelines that trigger on both schedules and performance thresholds keep models current without requiring manual intervention.
Question 38: Which of these is an unsupervised learning technique?
- Decision Trees
- Random Forest
- Support Vector Machines
- K-Means Clustering (Correct answer)
Correct answer: K-Means Clustering
K-Means Clustering is a prominent unsupervised learning technique used for partitioning a dataset into K distinct, non-overlapping subgroups or clusters. Unlike supervised methods, it does not require pre-labeled data; instead, it identifies inherent groupings based on the similarity of data points. This makes it ideal for tasks like customer segmentation or anomaly detection where labels are unknown.
Question 39: Which metric category is most critical when monitoring a deployed classification model in production?
- Number of API calls per second
- Model file size on disk
- Prediction confidence distribution and accuracy on live data (Correct answer)
- Training loss from the most recent training run
Correct answer: Prediction confidence distribution and accuracy on live data
Monitoring live prediction confidence and accuracy reveals model drift and performance degradation before it significantly impacts business outcomes.
Question 40: Which US federal agency primarily oversees algorithmic accountability in consumer financial products?
- CFPB (Correct answer)
- FTC
- SEC
- OCC
Correct answer: CFPB
The Consumer Financial Protection Bureau (CFPB) has primary authority over algorithmic models used in consumer lending, credit scoring, and financial products.
Question 41: Which NLP technique converts words into dense vector representations that capture semantic relationships?
- TF-IDF
- Bag of Words
- Word2Vec / Word Embeddings (Correct answer)
- Stemming
Correct answer: Word2Vec / Word Embeddings
Word embeddings like Word2Vec map words to dense vectors where semantically similar words have high cosine similarity.
Question 42: A stakeholder group has conflicting interpretations of model output. What is the most effective resolution approach?
- Choose the most senior stakeholder's interpretation automatically
- Convene a joint calibration session to standardize interpretation guidelines and document shared definitions (Correct answer)
- Let each group use their own interpretation independently
- Remove the conflicting output from the model entirely
Correct answer: Convene a joint calibration session to standardize interpretation guidelines and document shared definitions
A calibration session with documented shared definitions eliminates ambiguity and ensures consistent model output interpretation across teams.
Question 43: Which metric is commonly used to evaluate object detection models by measuring overlap between predicted and ground-truth bounding boxes?
- ROUGE Score
- BLEU Score
- Intersection over Union (IoU) (Correct answer)
- F1 Score
Correct answer: Intersection over Union (IoU)
IoU divides the area of overlap between the predicted and actual bounding boxes by their combined union area, ranging from 0 to 1.
Question 44: In AML practice, what is evidence-based practice?
- Using only the newest methods regardless of evidence
- Following only personal experience and intuition
- Integrating the best available research evidence with professional expertise and client needs (Correct answer)
- Doing whatever the client requests
Correct answer: Integrating the best available research evidence with professional expertise and client needs
Evidence-based practice combines rigorous research evidence, professional clinical expertise, and client/patient preferences and values to make informed decisions that optimize outcomes.
Question 45: In cohort analysis, what does retention rate measure?
- The average revenue per user cohort
- The churn rate of the bottom decile
- The percentage of new users acquired in a period
- The fraction of users from an initial cohort who remain active in a later period (Correct answer)
Correct answer: The fraction of users from an initial cohort who remain active in a later period
Retention rate tracks what fraction of users who joined in a given cohort period are still active N periods later.
Question 46: What is A/B testing in the context of ML model deployment?
- Running automated unit tests on model training code
- Comparing two model versions by routing live traffic to each and measuring performance differences (Correct answer)
- Testing models on two separate servers simultaneously
- Testing two different datasets against the same model
Correct answer: Comparing two model versions by routing live traffic to each and measuring performance differences
A/B testing routes a portion of real production traffic to a challenger model while the rest serves the incumbent, enabling data-driven model selection.
Question 47: A bank's loan model is audited and found to have higher false positive rates for minority applicants (incorrectly flagging them as high-risk). Which remediation technique directly targets this disparity?
- Increasing training data volume uniformly
- Applying reweighting or post-processing threshold adjustments per group (Correct answer)
- Adding regularization to reduce overfitting
- Switching to a more complex model architecture
Correct answer: Applying reweighting or post-processing threshold adjustments per group
Reweighting training samples or applying group-specific decision thresholds via post-processing are standard fairness interventions that directly correct disparate false positive rates.
Question 48: Which technique addresses class imbalance during model evaluation by computing metrics on a resampled dataset that reflects equal class distribution?
- Stratified k-fold cross-validation
- Balanced accuracy metric on original data (Correct answer)
- Bootstrap resampling with replacement
- SMOTE oversampling followed by standard evaluation
Correct answer: Balanced accuracy metric on original data
Balanced accuracy averages recall across classes, directly accounting for imbalance without resampling the dataset itself.
Question 49: What is MLOps in the context of advanced machine learning?
- A dataset versioning tool
- A programming language for machine learning
- The practice of combining ML development with operations to deploy and maintain models in production (Correct answer)
- A type of neural network architecture
Correct answer: The practice of combining ML development with operations to deploy and maintain models in production
MLOps applies DevOps principles to machine learning, automating training, deployment, monitoring, and retraining workflows.
Question 50: In the context of learning curves, what does a large gap between training and validation loss with both curves converging indicate?
- Underfitting requiring a more complex model
- High variance (overfitting) needing more regularization or data (Correct answer)
- Label noise causing inconsistent gradients
- Optimal model capacity with good generalization
Correct answer: High variance (overfitting) needing more regularization or data
A large persistent gap between training and validation loss with convergence indicates the model has memorized training data and is not generalizing well (high variance/overfitting).
Question 51: Which pandas function is most efficient for computing group-level summary statistics across a large DataFrame?
- df.groupby().agg() (Correct answer)
- df.apply() with a lambda
- df.merge() with a lookup table
- df.iterrows() with accumulation
Correct answer: df.groupby().agg()
groupby().agg() leverages vectorized operations internally, making it far more efficient than row-wise apply or iterrows for group aggregations.
Question 52: What is the purpose of max pooling in a convolutional neural network?
- Introduce non-linearity into the network
- Normalize activations across the batch
- Increase the spatial resolution of feature maps
- Reduce spatial dimensions while retaining dominant features (Correct answer)
Correct answer: Reduce spatial dimensions while retaining dominant features
Max pooling downsamples feature maps by taking the maximum value in each pooling window, reducing size while preserving strong activations.
Question 53: What does professional competency require of a AML practitioner?
- Maintaining current knowledge through continuing education and practicing only within areas of qualification (Correct answer)
- Accepting all work regardless of qualifications
- Learning only during formal schooling
- Relying solely on initial certification training
Correct answer: Maintaining current knowledge through continuing education and practicing only within areas of qualification
Professional competency requires practitioners to maintain current knowledge through ongoing education, stay informed about developments in their field, and only practice within the boundaries of their demonstrated competence.
Question 54: What does TF-IDF stand for in text feature extraction?
- Token Filter–Index Document Format
- Text Format–Inverse Data Frame
- Term Frequency–Inverse Document Frequency (Correct answer)
- Text Frequency–Inverse Document Function
Correct answer: Term Frequency–Inverse Document Frequency
TF-IDF weighs each word by how often it appears in a document (TF) penalized by how common it is across all documents (IDF).
Question 55: In AML practice, what is evidence-based practice?
- Doing whatever the client requests
- Following only personal experience and intuition
- Using only the newest methods regardless of evidence
- Integrating the best available research evidence with professional expertise and client needs (Correct answer)
Correct answer: Integrating the best available research evidence with professional expertise and client needs
Evidence-based practice combines rigorous research evidence, professional clinical expertise, and client/patient preferences and values to make informed decisions that optimize outcomes.
Question 56: In the context of LLM inference optimization, what does 'KV cache' refer to?
- Cached key-value attention matrices reused across autoregressive generation steps (Correct answer)
- GPU L2 cache for matrix multiplication kernels
- A key-value store for experiment metadata
- A Redis cache storing tokenized prompts
Correct answer: Cached key-value attention matrices reused across autoregressive generation steps
The KV cache stores computed key and value tensors from previous tokens during autoregressive decoding, avoiding redundant recomputation and dramatically accelerating text generation.
Question 57: What does model quantization achieve during ML model deployment?
- It encrypts model weights for secure deployment in regulated environments
- It increases model accuracy by adding more learnable parameters
- It reduces model size and speeds up inference by using lower-precision numerical representations (Correct answer)
- It converts regression models into classification models automatically
Correct answer: It reduces model size and speeds up inference by using lower-precision numerical representations
Quantization converts 32-bit floating point weights to 8-bit integers, drastically reducing model size and inference latency with minimal accuracy loss.
Question 58: What is the key difference between normalization and standardization?
- Normalization removes outliers; standardization does not
- They are identical techniques with different names
- Normalization encodes categories; standardization scales numerics
- Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance (Correct answer)
Correct answer: Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance
Normalization maps values to a fixed range like [0,1] while standardization transforms data to have mean 0 and standard deviation 1.
Question 59: What is the attention mechanism in transformer models primarily designed to do?
- Allow the model to weigh the relevance of different input tokens when producing each output (Correct answer)
- Reduce gradient vanishing during backpropagation
- Replace the need for positional encodings
- Increase the batch size during training
Correct answer: Allow the model to weigh the relevance of different input tokens when producing each output
Attention computes a weighted sum of value vectors, where weights reflect how relevant each input token is to the current output position.
Question 60: In Principal Component Analysis (PCA), what do the eigenvalues of the covariance matrix represent?
- The proportion of variance explained by each principal component (Correct answer)
- The number of clusters present in the data
- The correlation coefficients between original features
- The directions of maximum variance (principal components)
Correct answer: The proportion of variance explained by each principal component
Each eigenvalue quantifies the amount of variance in the data that is captured by its corresponding eigenvector (principal component direction).
Question 61: Which evaluation metric for generative language models measures how well a probability model predicts a sample by computing the exponential of average negative log-likelihood?
- ROC-AUC
- F1 Score
- BLEU
- Perplexity (Correct answer)
Correct answer: Perplexity
Perplexity quantifies how surprised a language model is by new text; lower perplexity indicates better predictive performance.
Question 62: Which regularization technique used in neural networks randomly sets a fraction of neuron activations to zero during training?
- Dropout (Correct answer)
- Early Stopping
- Batch Normalization
- L2 Weight Decay
Correct answer: Dropout
Dropout randomly zeroes a proportion of neuron outputs each forward pass, preventing co-adaptation of neurons and acting as an ensemble of many subnetworks.
Question 63: What does SHAP (SHapley Additive exPlanations) provide in ML reporting?
- A regularization technique for neural networks
- A method to handle missing data
- Global feature importances only
- Per-prediction, feature-level attributions explaining individual model outputs (Correct answer)
Correct answer: Per-prediction, feature-level attributions explaining individual model outputs
SHAP values assign each feature a contribution to a specific prediction, enabling both local (per-instance) and global interpretability.
Question 64: In computer vision, what is the role of a convolutional layer in a CNN?
- Apply learnable filters to detect local spatial features (Correct answer)
- Flatten the image into a 1D vector
- Normalize pixel values across channels
- Perform non-maximum suppression on detected objects
Correct answer: Apply learnable filters to detect local spatial features
Convolutional layers slide learnable filters across the input image to detect local patterns like edges, textures, and shapes.
Question 65: What does professional liability insurance protect a AML practitioner against?
- Employee health expenses
- Theft of office equipment
- Natural disaster damage to the office
- Financial loss from claims of negligence, errors, or omissions in professional services (Correct answer)
Correct answer: Financial loss from claims of negligence, errors, or omissions in professional services
Professional liability insurance (errors and omissions coverage) protects practitioners from the financial consequences of claims alleging negligence, mistakes, or failure to perform professional duties, covering legal defense costs and settlements.
Question 66: What does a DET (Detection Error Tradeoff) curve plot that distinguishes it from an ROC curve?
- Precision vs. recall on a linear scale
- True Positive Rate vs. False Positive Rate on log scale
- Lift vs. cumulative population percentage
- False Negative Rate vs. False Positive Rate on normal deviate axes (Correct answer)
Correct answer: False Negative Rate vs. False Positive Rate on normal deviate axes
DET curves plot FNR vs. FPR using normal deviate (probit) scales, which linearize performance curves for easier comparison of classifiers.
Question 67: In gradient boosting, what is the role of the shrinkage parameter (learning rate)?
- It determines the fraction of features sampled per tree
- It controls the depth of individual decision trees
- It scales each new tree's contribution, reducing overfitting at the cost of requiring more trees (Correct answer)
- It sets the minimum number of samples required for a split
Correct answer: It scales each new tree's contribution, reducing overfitting at the cost of requiring more trees
The shrinkage/learning rate multiplies each tree's output before adding it to the ensemble, making each step conservative and reducing overfitting while necessitating more iterations.
Question 68: What is canary deployment in ML production systems?
- Using a lightweight fallback model when the primary model fails
- Deploying a model only for internal QA testing
- Gradually rolling out a new model to a small subset of users before a full release (Correct answer)
- Testing new models in a completely isolated validation environment
Correct answer: Gradually rolling out a new model to a small subset of users before a full release
Canary deployment incrementally increases traffic to a new model version, allowing real-world validation with minimal risk exposure.
Question 69: In MLflow, what is the purpose of the 'Model Registry' component?
- To store raw training datasets
- To schedule automated retraining pipelines
- To manage model versioning, staging, and production transitions (Correct answer)
- To visualize training metrics in real time
Correct answer: To manage model versioning, staging, and production transitions
MLflow Model Registry provides a centralized hub for managing the full lifecycle of ML models including versioning and stage transitions (Staging, Production, Archived).
Question 70: Which NLP preprocessing step reduces words to their base or root form by removing suffixes (e.g., 'running' → 'run')?
- Stemming (Correct answer)
- POS tagging
- Lemmatization
- Dependency parsing
Correct answer: Stemming
Stemming applies rule-based suffix stripping to reduce words to an approximate root, which may not be a valid dictionary word.
Question 71: Which dimensionality reduction technique projects data onto directions of maximum variance?
- UMAP
- t-SNE
- PCA (Correct answer)
- LDA
Correct answer: PCA
PCA (Principal Component Analysis) finds orthogonal axes of maximum variance to compress feature dimensions.
Question 72: What is a REST API in the context of machine learning model deployment?
- A version control system for datasets
- A database for storing model artifacts
- A standardized interface that allows applications to send data to a model and receive predictions via HTTP (Correct answer)
- A monitoring dashboard for deployed ML models
Correct answer: A standardized interface that allows applications to send data to a model and receive predictions via HTTP
A REST API exposes model inference as HTTP endpoints so any downstream application can request predictions via standard web requests.
Question 73: What is the computational complexity of a single k-NN prediction for a dataset with N training samples and D features?
- O(N * D) (Correct answer)
- O(D)
- O(log N)
- O(N^2 * D)
Correct answer: O(N * D)
A brute-force k-NN prediction requires computing the distance from the query point to all N training points, each requiring O(D) operations, giving O(N * D) total.
Question 74: What does the Bellman equation fundamentally express in reinforcement learning?
- The recursive relationship between a state's value and the values of its successor states (Correct answer)
- The gradient of the policy loss with respect to network parameters
- The optimal epsilon decay schedule for exploration
- The backpropagation update rule for deep Q-networks
Correct answer: The recursive relationship between a state's value and the values of its successor states
The Bellman equation expresses value recursively: the value of a state equals the immediate reward plus the discounted value of successor states, enabling dynamic programming solutions.
Question 75: A compliance officer requests interpretability reports for a credit scoring model. Which approach best meets both technical and regulatory communication needs?
- Provide only the model's source code
- Share only the model's overall accuracy metric
- Claim the model is a black box and cannot be explained
- Generate SHAP-based feature importance reports translated into plain-language decision explanations aligned with regulatory requirements (Correct answer)
Correct answer: Generate SHAP-based feature importance reports translated into plain-language decision explanations aligned with regulatory requirements
SHAP-based explanations translated into plain language satisfy both technical interpretability needs and regulatory requirements for decision transparency.
Question 76: Which transformer-based model introduced bidirectional context for language understanding and set new NLP benchmarks in 2018?
- BERT (Correct answer)
- ELMo
- Seq2Seq
- GPT-2
Correct answer: BERT
BERT (Bidirectional Encoder Representations from Transformers) pre-trains on masked language modeling using full left and right context simultaneously.
Question 77: Which cross-validation strategy is most appropriate when your dataset has significant temporal ordering?
- Shuffle-split CV
- Stratified k-fold
- Leave-one-out CV
- Time-series split (walk-forward) (Correct answer)
Correct answer: Time-series split (walk-forward)
Time-series split (walk-forward validation) prevents data leakage by always training on past data and validating on future data.
Question 78: In a continuous integration pipeline for ML models, a 'regression gate' is best defined as:
- An automated check that blocks deployment if a new model underperforms the current production model on a reference dataset (Correct answer)
- A holdout set used exclusively for regression task evaluation
- A version control branch strategy for model artifacts
- A regularization layer added before model deployment
Correct answer: An automated check that blocks deployment if a new model underperforms the current production model on a reference dataset
A regression gate enforces a minimum quality bar by comparing a candidate model against the current champion before allowing promotion to production.
Question 79: What does CI/CD stand for in an MLOps pipeline?
- Central Intelligence / Cloud Deployment
- Containerized Inference / Continuous Development
- Continuous Integration / Continuous Deployment (Correct answer)
- Continuous Improvement / Continuous Diagnostics
Correct answer: Continuous Integration / Continuous Deployment
CI/CD automates testing and deployment pipelines so new model versions can be validated and released quickly and reliably.
Question 80: What is the purpose of peer review in AML professional practice?
- To evaluate work quality through assessment by qualified colleagues and promote continuous improvement (Correct answer)
- To reduce workload through delegation
- To create competition between colleagues
- To determine who should be promoted
Correct answer: To evaluate work quality through assessment by qualified colleagues and promote continuous improvement
Peer review provides objective quality assessment by qualified professionals, identifying areas for improvement, validating practices, and promoting professional accountability and continuous quality enhancement.
Question 81: What is the primary purpose of model explainability tools like SHAP in production ML systems?
- To automate feature engineering pipelines
- To compress model artifacts for faster deployment
- To speed up model inference at scale
- To explain individual predictions and ensure model transparency for stakeholders and auditors (Correct answer)
Correct answer: To explain individual predictions and ensure model transparency for stakeholders and auditors
SHAP attributes each feature's contribution to individual predictions, supporting transparency and regulatory compliance in production.
Question 82: When planning compute resources for distributed ML training, what is 'communication overhead' and why does it matter?
- Latency in reading training data from disk on a single node
- The cost of sending emails between team members
- Time engineers spend in meetings, which reduces coding productivity
- The time gradient synchronization across workers takes, which can dominate wall-clock training time when the model-to-data ratio is high (Correct answer)
Correct answer: The time gradient synchronization across workers takes, which can dominate wall-clock training time when the model-to-data ratio is high
In distributed training, gradient synchronization overhead can exceed compute time per step, especially for large models on slow interconnects, making network topology a critical planning factor.
Question 83: What is a machine learning data pipeline?
- A method for hyperparameter tuning
- A storage system for training datasets
- An automated sequence of data processing steps from ingestion to model-ready format (Correct answer)
- A visualization tool for feature distributions
Correct answer: An automated sequence of data processing steps from ingestion to model-ready format
A data pipeline automates and chains data collection, cleaning, transformation, and feature engineering steps for reproducibility.
Question 84: What is the primary purpose of the 'kernel trick' in spectral clustering?
- To select the optimal number of clusters automatically
- To normalize the graph Laplacian before clustering
- To implicitly compute similarities in high-dimensional feature spaces without explicit mapping (Correct answer)
- To speed up eigendecomposition of the similarity matrix
Correct answer: To implicitly compute similarities in high-dimensional feature spaces without explicit mapping
In spectral clustering, kernel functions allow computation of pairwise similarities in a high-dimensional space via the kernel trick, enabling non-linear cluster discovery.
Question 85: What is the key purpose of a 'threshold analysis' in ML risk management for classification models?
- Measuring the minimum dataset size needed for statistical significance
- Evaluating how different decision thresholds trade off false positives and false negatives to match risk tolerance (Correct answer)
- Setting the maximum allowed training epochs to prevent overfitting
- Determining the optimal number of hidden layers
Correct answer: Evaluating how different decision thresholds trade off false positives and false negatives to match risk tolerance
Threshold analysis explores the precision-recall or FPR-TPR trade-off across all classification cut-points so risk managers can select the threshold that aligns with acceptable error costs.
Question 86: Under the NIST AI Risk Management Framework (AI RMF), the 'Govern' function primarily addresses:
- Organizational policies, accountability structures, and culture for AI risk (Correct answer)
- Real-time model monitoring pipelines
- Technical performance benchmarking
- Dataset curation and labeling standards
Correct answer: Organizational policies, accountability structures, and culture for AI risk
The Govern function in NIST AI RMF establishes organizational context, accountability, and culture that enable effective AI risk management across the other functions.
Question 87: Which principle in the OECD AI Principles (2019) most directly addresses the obligation to provide recourse when AI systems cause harm?
- Robustness, security, and safety
- Transparency and explainability
- Accountability (Correct answer)
- Inclusive growth and sustainable development
Correct answer: Accountability
The OECD Accountability principle requires AI actors to be responsible for the proper functioning of AI systems and for addressing any harm they cause.
Question 88: Which unsupervised algorithm is best suited for detecting arbitrarily shaped clusters in noisy data?
- K-Means
- DBSCAN (Correct answer)
- Hierarchical Agglomerative Clustering
- Gaussian Mixture Models
Correct answer: DBSCAN
DBSCAN (Density-Based Spatial Clustering of Applications with Noise) discovers clusters of arbitrary shape and explicitly labels outliers as noise points.
Question 89: What is the role of a 'vector database' (e.g., Pinecone, Weaviate, Chroma) in modern ML systems?
- Scheduling distributed training jobs
- Logging experiment metrics and parameters
- Storing and versioning trained model weights
- Efficient similarity search over high-dimensional embedding vectors (Correct answer)
Correct answer: Efficient similarity search over high-dimensional embedding vectors
Vector databases index high-dimensional embeddings and support approximate nearest-neighbor (ANN) search, enabling fast semantic retrieval for RAG and recommendation systems.
Question 90: Which optimizer adapts per-parameter learning rates using exponentially weighted moving averages of both gradients and their squares?
- RMSprop
- AdaGrad
- Adam (Correct answer)
- SGD with momentum
Correct answer: Adam
Adam combines first-moment (gradient) and second-moment (squared gradient) exponential moving averages with bias correction, making it more adaptive than RMSprop or AdaGrad alone.
Question 91: What is the goal of clustering in unsupervised learning?
- Encoding categorical variables
- Minimizing data quality
- Classifying known labels
- Grouping similar data points (Correct answer)
Correct answer: Grouping similar data points
The primary goal of clustering in unsupervised learning is to discover inherent structures within unlabeled datasets by grouping similar data points together. Algorithms achieve this by identifying patterns and relationships that allow them to form distinct clusters, where data points within a cluster are more similar to each other than to those in other clusters. This process helps in uncovering hidden categories or segments within the data.
Question 92: Which statistical test is most appropriate for detecting covariate shift between training and production data distributions?
- Two-sample Kolmogorov-Smirnov test (Correct answer)
- Paired t-test
- Chi-squared goodness-of-fit test
- ANOVA
Correct answer: Two-sample Kolmogorov-Smirnov test
The two-sample KS test compares the empirical CDFs of two continuous distributions without assuming normality, making it ideal for detecting feature drift.
Question 93: In AML practice, what is the purpose of a standard operating procedure (SOP)?
- To document step-by-step instructions for routine tasks to ensure consistency and quality (Correct answer)
- To satisfy management preferences only
- To restrict employee creativity
- To create unnecessary paperwork
Correct answer: To document step-by-step instructions for routine tasks to ensure consistency and quality
SOPs provide standardized, detailed instructions for routine operations, ensuring consistency, quality, efficiency, and safety regardless of which qualified individual performs the task.
Question 94: What is the purpose of temperature scaling in post-hoc model calibration?
- It adjusts the softmax temperature using a single learned scalar to align predicted probabilities with empirical frequencies (Correct answer)
- It applies label smoothing retroactively to training labels
- It ensembles multiple models at different temperatures
- It reduces inference latency by quantizing model weights
Correct answer: It adjusts the softmax temperature using a single learned scalar to align predicted probabilities with empirical frequencies
Temperature scaling divides logits by a learned temperature T before softmax, uniformly adjusting confidence levels to better match empirical accuracy without changing predictions.
Question 95: What is the purpose of cross-validation in model evaluation reporting?
- To tune hyperparameters faster
- To increase training data size
- To provide an unbiased estimate of model performance on unseen data (Correct answer)
- To detect data leakage
Correct answer: To provide an unbiased estimate of model performance on unseen data
Cross-validation estimates generalization performance by rotating held-out folds, reducing reliance on a single train-test split that may be lucky or unlucky.
Question 96: What is training-serving skew in machine learning deployment?
- Discrepancies between how features are computed during offline training versus online model serving (Correct answer)
- The time lag between completing model training and completing deployment
- The accuracy gap between training and test set performance
- The mismatch between model size and available hardware capacity
Correct answer: Discrepancies between how features are computed during offline training versus online model serving
Training-serving skew occurs when feature computation logic differs between offline training and online serving, causing unexpected prediction errors in production.
Question 97: A model's calibration is assessed using reliability diagrams. A poorly calibrated model that consistently predicts 80% confidence but achieves only 60% actual accuracy is said to be:
- Exhibiting concept drift
- Biased toward the majority class
- Underconfident
- Overconfident (Correct answer)
Correct answer: Overconfident
Overconfidence means predicted probabilities are systematically higher than the actual observed frequencies, which reliability diagrams reveal as the calibration curve falling below the diagonal.
Question 98: Which encoding technique assigns integer values to categories that have a natural order?
- Ordinal encoding (Correct answer)
- Binary encoding
- Frequency encoding
- One-hot encoding
Correct answer: Ordinal encoding
Ordinal encoding maps ordered categories such as low/medium/high to integers that preserve their natural rank.
Advanced Machine Learning (AML) Certification
The AML certification validates advanced competency across the full machine learning lifecycle, covering core algorithms, feature engineering, model deployment, NLP, computer vision, and professional skills including AI ethics, governance, and stakeholder communication.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds