Mixed Deck — All AML Topics Flashcards
100 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 20 Mixed Deck — All AML Topics flashcards as text
Which of these is an unsupervised learning technique?
Answer: K-Means Clustering
K-Means Clustering is a prominent unsupervised learning technique used for partitioning a dataset into K distinct, non-overlapping subgroups or clusters. Unlike supervised methods, it does not require pre-labeled data; instead, it identifies inherent groupings based on the similarity of data points. This makes it ideal for tasks like customer segmentation or anomaly detection where labels are unknown.
Which statement correctly describes the bias-variance tradeoff in the context of k-Nearest Neighbors (kNN)?
Answer: Larger k increases bias and decreases variance
A larger k smooths the decision boundary by averaging more neighbors, increasing bias (underfitting) but reducing variance (sensitivity to noise).
In the context of LLM inference optimization, what does 'KV cache' refer to?
Answer: Cached key-value attention matrices reused across autoregressive generation steps
The KV cache stores computed key and value tensors from previous tokens during autoregressive decoding, avoiding redundant recomputation and dramatically accelerating text generation.
What is the goal of data governance?
Answer: To manage data integrity and compliance
The goal of data governance is to establish and enforce policies, procedures, and standards for managing an organization's data assets. This ensures data integrity, quality, security, and compliance with relevant regulations and ethical guidelines. Effective data governance is essential for reliable AI systems, as it guarantees that the data used for training and operation is trustworthy and legally sound.
Which technique is most appropriate for communicating model uncertainty to executive stakeholders?
Answer: Present confidence intervals alongside predictions with plain-language explanations
Confidence intervals paired with plain-language explanations help executives understand prediction reliability without overwhelming technical detail.
Which linkage method in hierarchical clustering is most sensitive to outliers?
Answer: Single linkage
Single linkage defines cluster distance as the minimum distance between any two points across clusters, making it highly sensitive to outliers that can form long 'chaining' clusters.
A feature importance analysis reveals that a lending model heavily weights 'neighborhood' as a proxy feature. This MOST directly suggests the presence of:
Answer: Proxy discrimination using a legally protected attribute
Neighborhood is a well-known proxy for race and national origin; heavy reliance on it constitutes proxy discrimination even if protected attributes are excluded.
Which of the following BEST illustrates a conflict of interest for an ML practitioner?
Answer: Evaluating a model built by a vendor in which the practitioner holds equity
Holding financial interest in a vendor while objectively evaluating that vendor's product creates a material conflict of interest.
When would you prefer Naive Bayes over Logistic Regression for text classification?
Answer: When training data is scarce and the conditional independence assumption is approximately met
Naive Bayes performs well with small datasets because it estimates feature probabilities independently, requiring far fewer parameters than logistic regression to learn.
Which metric category is most critical when monitoring a deployed classification model in production?
Answer: Prediction confidence distribution and accuracy on live data
Monitoring live prediction confidence and accuracy reveals model drift and performance degradation before it significantly impacts business outcomes.
Which technique addresses class imbalance during model evaluation by computing metrics on a resampled dataset that reflects equal class distribution?
Answer: Balanced accuracy metric on original data
Balanced accuracy averages recall across classes, directly accounting for imbalance without resampling the dataset itself.
What is the purpose of cross-validation in model evaluation reporting?
Answer: To provide an unbiased estimate of model performance on unseen data
Cross-validation estimates generalization performance by rotating held-out folds, reducing reliance on a single train-test split that may be lucky or unlucky.
In federated learning, which privacy concern is MOST directly mitigated compared to centralized training?
Answer: Raw training data never leaving local devices
Federated learning's core privacy property is that raw data remains on-device; only model updates (gradients or weights) are shared with the central aggregator.
In AML practice, what is evidence-based practice?
Answer: Integrating the best available research evidence with professional expertise and client needs
Evidence-based practice combines rigorous research evidence, professional clinical expertise, and client/patient preferences and values to make informed decisions that optimize outcomes.
A compliance officer requests interpretability reports for a credit scoring model. Which approach best meets both technical and regulatory communication needs?
Answer: Generate SHAP-based feature importance reports translated into plain-language decision explanations aligned with regulatory requirements
SHAP-based explanations translated into plain language satisfy both technical interpretability needs and regulatory requirements for decision transparency.
Which of the following best describes 'regulatory sandboxes' in the context of AI governance?
Answer: Controlled frameworks allowing companies to test AI products under relaxed rules with regulatory oversight
Regulatory sandboxes let companies pilot innovative AI systems in live environments under temporary regulatory relaxation and direct supervisory oversight.
What does a stochastic policy output for a given state in reinforcement learning?
Answer: A probability distribution over all possible actions
A stochastic policy maps each state to a probability distribution over actions, enabling exploration by sampling from this distribution during training.
When evaluating a regression model, which metric is scale-independent and allows comparison across datasets with different target variable magnitudes?
Answer: Mean Absolute Percentage Error (MAPE)
MAPE expresses errors as a percentage of actual values, making it scale-independent and comparable across datasets with different target magnitudes, unlike absolute metrics.
Which technique helps detect if a model has learned spurious correlations by testing it on out-of-distribution examples with controlled variations?
Answer: Behavioral testing with invariance and directional expectation tests (CheckList)
CheckList-style behavioral testing defines expected model behaviors under transformations (invariance) and targeted perturbations (direction tests) to expose learned shortcuts.
Which principle in the OECD AI Principles (2019) most directly addresses the obligation to provide recourse when AI systems cause harm?
Answer: Accountability
The OECD Accountability principle requires AI actors to be responsible for the proper functioning of AI systems and for addressing any harm they cause.