โ† All Data Science Flashcard Decks

Supervised Learning Algorithms Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Supervised Learning Algorithms flashcards as text
  1. In a decision tree, what does a higher Gini impurity at a node indicate?

    Answer: The classes are more mixed at that node

    Gini impurity increases as class distribution within a node becomes more mixed.

  2. Which supervised algorithm is most prone to overfitting if grown without depth limits?

    Answer: An unpruned decision tree

    A fully grown unpruned tree can memorize training data, causing overfitting.

  3. What is the primary purpose of the kernel trick in support vector machines?

    Answer: To enable linear separation in a higher-dimensional space without explicit mapping

    The kernel trick computes inner products in a higher-dimensional space implicitly, allowing nonlinear separation.

  4. In logistic regression, what does the sigmoid function output represent?

    Answer: A probability between 0 and 1

    The sigmoid maps the linear combination of inputs to a probability between 0 and 1.

  5. Why does k-nearest neighbors require feature scaling?

    Answer: Distance calculations are dominated by large-range features

    Unscaled features with larger ranges disproportionately influence the distance metric.

  6. What does the 'naive' assumption in Naive Bayes refer to?

    Answer: Features are conditionally independent given the class

    Naive Bayes assumes features are conditionally independent given the class label.

  7. Which metric is most appropriate for evaluating a classifier on a heavily imbalanced dataset?

    Answer: F1 score or AUC-PR

    F1 score or precision-recall AUC better reflect performance when classes are imbalanced than accuracy.