Supervised Learning Algorithms Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised Learning Algorithms flashcards as text
In random forests, what technique reduces correlation between individual trees?
Answer: Random feature subsampling at each split
Selecting a random subset of features at each split decorrelates the trees in a random forest.
What is the key difference between bagging and boosting?
Answer: Bagging trains models independently in parallel; boosting trains sequentially focusing on errors
Bagging builds independent models in parallel while boosting builds models sequentially to correct prior errors.
In gradient boosting, what do successive trees primarily fit?
Answer: The residual errors of previous trees
Each new tree in gradient boosting fits the residuals (errors) left by the prior ensemble.
What does a large C value in an SVM control?
Answer: A narrower margin penalizing misclassifications more
A large C penalizes misclassifications heavily, producing a narrower margin and less regularization.
Which loss function is typically minimized in standard linear regression?
Answer: Mean squared error
Ordinary least squares linear regression minimizes the mean squared error between predictions and targets.
What problem does L1 (Lasso) regularization help address that L2 does not?
Answer: It can drive some coefficients exactly to zero, performing feature selection
L1 regularization can shrink coefficients to exactly zero, effectively selecting features.
In a confusion matrix, what does recall measure?
Answer: True positives over all actual positives
Recall is the proportion of actual positives correctly identified (TP / (TP + FN)).