Supervised & Unsupervised Learning Algorithms Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised & Unsupervised Learning Algorithms flashcards as text
In a gradient boosting model, what does the 'learning rate' (shrinkage) parameter control?
Answer: The fraction by which each tree's contribution is scaled down
The learning rate (shrinkage) multiplies each new tree's contribution, preventing overfitting by making the ensemble learn more slowly and conservatively.
Which unsupervised algorithm is best suited for detecting arbitrarily shaped clusters in noisy data?
Answer: DBSCAN
DBSCAN (Density-Based Spatial Clustering of Applications with Noise) discovers clusters of arbitrary shape and explicitly labels outliers as noise points.
What is the primary purpose of using kernel functions in Support Vector Machines?
Answer: To implicitly map data into a higher-dimensional feature space where it becomes linearly separable
Kernel functions compute dot products in a higher-dimensional space without explicitly transforming the data, enabling SVMs to find non-linear decision boundaries.
In Expectation-Maximization (EM) for Gaussian Mixture Models, what happens during the E-step?
Answer: Soft cluster membership probabilities are computed for each data point
The E-step computes the posterior probability (responsibility) that each Gaussian component generated each data point, given the current parameter estimates.
Which statement correctly describes the bias-variance tradeoff in the context of k-Nearest Neighbors (kNN)?
Answer: Larger k increases bias and decreases variance
A larger k smooths the decision boundary by averaging more neighbors, increasing bias (underfitting) but reducing variance (sensitivity to noise).
What does the 'elbow method' evaluate when selecting the optimal number of clusters for K-Means?
Answer: Within-cluster sum of squares (inertia) as a function of k
The elbow method plots within-cluster sum of squares (inertia) against k and selects the point where the rate of decrease sharply diminishes, forming an 'elbow.'
In a Random Forest, how does the 'max_features' hyperparameter reduce correlation between trees?
Answer: It restricts each split to consider only a random subset of features
By considering only a random subset of features at each split, trees are forced to use different predictors, reducing correlation and improving ensemble diversity.