Machine Learning Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Machine Learning flashcards as text
What does the VC dimension measure in statistical learning theory?
Answer: The capacity or complexity of a hypothesis class, defined by the largest set it can shatter
The VC dimension is the size of the largest set of points that a hypothesis class can classify correctly in all possible ways, quantifying model capacity.
In principal component analysis (PCA), what do the principal components represent?
Answer: Orthogonal directions of maximum variance in the data
PCA finds orthogonal axes (principal components) that capture the directions of maximum variance, sorted from highest to lowest explained variance.
Which cross-validation strategy is most appropriate when the dataset has a temporal ordering?
Answer: Time-series split (walk-forward validation)
Time-series split trains on past observations and validates on future ones, preserving temporal order and preventing data leakage from future to past.
In the Expectation-Maximization (EM) algorithm for Gaussian Mixture Models, what happens in the E-step?
Answer: Soft responsibilities (posterior probabilities of cluster membership) are computed for each point
The E-step computes the expected cluster membership (responsibility) of each data point given the current model parameters using Bayes' theorem.
What does SMOTE (Synthetic Minority Oversampling TEchnique) do to address class imbalance?
Answer: Creates synthetic minority samples by interpolating between existing minority instances
SMOTE generates new synthetic minority samples by selecting a minority point and interpolating along the line segment to one of its k nearest minority neighbors.
Which metric is most informative when evaluating a classifier on a highly imbalanced dataset where false negatives are costly?
Answer: Recall (Sensitivity)
Recall measures the proportion of actual positives correctly identified; when false negatives are costly and classes are imbalanced, accuracy is misleading and recall directly captures missed positives.
In a neural network, what is the vanishing gradient problem?
Answer: Gradients become extremely small in early layers, slowing or halting learning
During backpropagation through many layers with saturating activations (e.g., sigmoid), gradients are repeatedly multiplied by small values, shrinking exponentially and causing early layers to learn very slowly.