Supervised Learning Algorithms Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised Learning Algorithms flashcards as text
What distinguishes Ridge regression from ordinary least squares?
Answer: Ridge adds an L2 penalty that shrinks coefficients toward zero without eliminating them
Ridge regression (L2 regularization) adds a penalty equal to the sum of squared coefficients, shrinking them uniformly toward zero but rarely setting any exactly to zero.
In a random forest, what is the purpose of feature randomness (selecting a subset of features at each split)?
Answer: It decorrelates the individual trees, reducing ensemble variance
By considering only a random subset of features at each split, trees in a random forest are de-correlated, so their errors are less likely to coincide, reducing overall variance.
Which scenario is most likely to benefit from using a polynomial regression model instead of linear regression?
Answer: When the relationship between the feature and target is curved or non-linear
Polynomial regression extends linear regression by adding polynomial terms, allowing it to fit curved relationships between predictors and a continuous outcome.
What is the purpose of the learning rate in gradient descent optimization?
Answer: It controls the step size taken in the direction of the negative gradient
The learning rate scales the gradient to determine how large a step is taken when updating parameters; too large causes divergence, too small causes slow convergence.
In logistic regression, what function maps the linear combination of inputs to a probability?
Answer: Sigmoid (logistic) function
The sigmoid function maps any real-valued linear combination to a value between 0 and 1, which can be interpreted as a class probability in binary logistic regression.
Which metric is most appropriate for evaluating a classifier on a highly imbalanced dataset where the minority class is critical?
Answer: F1 score
F1 score is the harmonic mean of precision and recall, making it sensitive to performance on the minority class and appropriate when false negatives and false positives both have high cost.
What is early stopping in neural network training?
Answer: Halting training when validation loss stops improving to prevent overfitting
Early stopping monitors validation performance and halts training when it stops improving, acting as a regularizer that prevents the model from overfitting the training set.