← All MS-DS Master of Data science Flashcard Decks

Supervised Learning Algorithms Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Supervised Learning Algorithms flashcards as text
  1. What is the primary advantage of gradient boosting over a single decision tree?

    Answer: It sequentially corrects errors of previous models, reducing bias

    Gradient boosting builds an ensemble by sequentially fitting new trees to the residual errors of the combined model so far, progressively reducing bias.

  2. In linear regression, what does the coefficient of determination (R²) measure?

    Answer: The proportion of variance in the target explained by the model

    R² measures the fraction of the total variance in the dependent variable that is captured by the model's predictions, ranging from 0 to 1.

  3. Which criterion is most commonly used to measure node impurity in classification trees?

    Answer: Gini impurity

    Gini impurity measures the probability of misclassifying a randomly chosen element and is the default splitting criterion in most classification tree implementations.

  4. A Naive Bayes classifier assumes which key property about features?

    Answer: Features are conditionally independent given the class label

    Naive Bayes applies Bayes' theorem with the 'naive' assumption that each feature is conditionally independent of every other feature given the class label.

  5. What is the effect of increasing the regularization strength (C parameter) in an SVM?

    Answer: The margin widens and misclassifications are penalized less

    In SVM, a smaller C allows more margin violations (softer margin), widening the margin and reducing sensitivity to individual training points.

  6. In the context of neural network training, what problem does the vanishing gradient address?

    Answer: Gradients become negligibly small in early layers, preventing effective weight updates

    Vanishing gradients occur when backpropagated gradients shrink exponentially through many layers, making early layers learn extremely slowly or not at all.

  7. Which ensemble method trains multiple models in parallel on random subsets of data and averages their predictions?

    Answer: Random Forest

    Random Forest builds multiple decision trees independently on bootstrapped data samples and random feature subsets, then aggregates their outputs via majority vote or averaging.