Knowledge Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Knowledge flashcards as text
Which algorithm is a non-parametric, instance-based learning method that classifies new points based on the majority label among their nearest neighbors?
Answer: K-Nearest Neighbors (KNN)
KNN stores training examples and classifies new instances by a vote among the k closest training points in feature space.
What is the difference between precision and recall in a classification context?
Answer: Precision is TP/(TP+FP); recall is TP/(TP+FN)
Precision measures how many predicted positives are actually positive, while recall measures how many actual positives were correctly identified.
Which data structure does a decision tree use to partition feature space?
Answer: Recursive binary splits based on feature thresholds
Decision trees recursively split the data using axis-aligned thresholds on individual features, creating a hierarchical partition of the input space.
What is 'feature engineering' in a data science workflow?
Answer: Creating, transforming, or selecting input variables to improve model performance
Feature engineering involves creating or transforming raw variables to make patterns more learnable by a model.
In a confusion matrix for binary classification, what does a False Negative (FN) represent?
Answer: Model predicted negative; actual class is positive
A False Negative occurs when the model predicts the negative class but the true label is positive, meaning a real positive was missed.
Which regularization technique adds the sum of absolute values of coefficients to the loss function?
Answer: Lasso (L1)
Lasso (L1) regularization penalizes the sum of absolute coefficient values, which can shrink some coefficients to exactly zero for feature selection.
What is the Central Limit Theorem's key implication for data science practice?
Answer: The sampling distribution of the mean approaches normality as sample size grows, regardless of the population distribution
The CLT states that sample means become approximately normally distributed with large enough samples, enabling parametric inference even for non-normal populations.