Machine Learning Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Machine Learning flashcards as text
What is the purpose of the learning rate in gradient descent optimization?
Answer: It scales the step size taken in the direction of the negative gradient
The learning rate is a scalar that multiplies the gradient to determine how large a step is taken toward the loss minimum during each parameter update.
Which concept describes the phenomenon where a feature's importance appears artificially low because its effect is shared with correlated features in a model?
Answer: Multicollinearity
Multicollinearity among predictors causes their individual coefficients or feature importances to be unstable and potentially misleading because the model cannot isolate their separate contributions.
In reinforcement learning, what is the Bellman equation used for?
Answer: Expressing the value of a state as the immediate reward plus discounted value of successor states
The Bellman equation decomposes the value function recursively: V(s) = E[r + γV(s')], forming the foundation of dynamic programming and Q-learning.
What distinguishes a generative model from a discriminative model?
Answer: Generative models learn the joint distribution P(X,Y); discriminative models learn P(Y|X) directly
Generative models model how data is generated (joint distribution), enabling them to synthesize new samples, while discriminative models focus solely on the decision boundary P(Y|X).
In XGBoost, what is the role of the 'lambda' (L2) regularization parameter?
Answer: It penalizes the squared magnitude of leaf weights to prevent overfitting
The lambda parameter in XGBoost adds an L2 penalty on leaf weight values in the objective function, shrinking them toward zero to reduce model complexity.
Which technique is used in neural architecture search and meta-learning to enable rapid adaptation to new tasks with few examples?
Answer: Model-Agnostic Meta-Learning (MAML)
MAML learns an initialization of model parameters that can be quickly fine-tuned to new tasks with just a few gradient steps and few labeled examples.
What is the curse of dimensionality's primary implication for distance-based machine learning algorithms like k-NN?
Answer: In high dimensions, all pairwise distances converge, making nearest-neighbor search unreliable
As dimensionality grows, the ratio of maximum to minimum pairwise distances approaches 1, so 'nearest' neighbors are nearly as far away as 'farthest' ones, undermining distance-based reasoning.