Supervised Learning Algorithms Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised Learning Algorithms flashcards as text
In a decision tree, what does a higher Gini impurity at a node indicate?
Answer: The classes are more mixed at that node
Gini impurity increases as class distribution within a node becomes more mixed.
Which supervised algorithm is most prone to overfitting if grown without depth limits?
Answer: An unpruned decision tree
A fully grown unpruned tree can memorize training data, causing overfitting.
What is the primary purpose of the kernel trick in support vector machines?
Answer: To enable linear separation in a higher-dimensional space without explicit mapping
The kernel trick computes inner products in a higher-dimensional space implicitly, allowing nonlinear separation.
In logistic regression, what does the sigmoid function output represent?
Answer: A probability between 0 and 1
The sigmoid maps the linear combination of inputs to a probability between 0 and 1.
Why does k-nearest neighbors require feature scaling?
Answer: Distance calculations are dominated by large-range features
Unscaled features with larger ranges disproportionately influence the distance metric.
What does the 'naive' assumption in Naive Bayes refer to?
Answer: Features are conditionally independent given the class
Naive Bayes assumes features are conditionally independent given the class label.
Which metric is most appropriate for evaluating a classifier on a heavily imbalanced dataset?
Answer: F1 score or AUC-PR
F1 score or precision-recall AUC better reflect performance when classes are imbalanced than accuracy.