โ† All Data Science Flashcard Decks

Data Science/Questions/Data Science 1 Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Science/Questions/Data Science 1 flashcards as text
  1. Which of the following best describes the bias-variance tradeoff in machine learning?

    Answer: A model with high bias underfits the data, while a model with high variance overfits it

    High bias means the model makes strong assumptions and fails to capture patterns (underfitting), while high variance means the model is too sensitive to training data (overfitting). The tradeoff involves finding the right complexity to minimize total error.

  2. When handling missing data in a dataset, which approach is most likely to introduce bias if the data is not missing at random (MNAR)?

    Answer: Listwise deletion (removing all rows with missing values)

    Listwise deletion removes entire rows with any missing value. When data is MNAR, the missingness is related to the unobserved value itself, so deletion systematically removes certain patterns and introduces significant bias into the remaining dataset.

  3. What is the main advantage of using gradient boosting over a single decision tree?

    Answer: Gradient boosting sequentially corrects the errors of prior trees, improving predictive accuracy

    Gradient boosting builds an ensemble of trees sequentially, where each new tree focuses on correcting the residual errors left by the previous trees. This iterative error correction typically yields much higher accuracy than any single decision tree.

  4. Which evaluation metric is most appropriate when the cost of false negatives is much higher than the cost of false positives, such as in cancer screening?

    Answer: Recall (Sensitivity)

    Recall measures the proportion of actual positives correctly identified (TP / (TP + FN)). When missing a true positive (false negative) is very costly, as in disease screening, maximizing recall ensures the fewest real cases are missed.

  5. In feature selection, what does the term 'multicollinearity' refer to and why is it a problem in linear regression?

    Answer: When independent features are highly correlated with each other, making coefficient estimates unreliable

    Multicollinearity occurs when independent variables are highly correlated with one another. In linear regression this inflates the variance of coefficient estimates, making them unstable and difficult to interpret, even though the overall model fit may appear adequate.

  6. Which statistical test is most appropriate for determining whether two independent groups have significantly different means, assuming normally distributed data and unknown but equal variances?

    Answer: Independent samples t-test (Student's t-test)

    The independent samples t-test (Student's t-test) is designed to compare the means of two separate, unrelated groups when the data is approximately normally distributed and the variances are assumed to be equal. It produces a t-statistic and p-value to assess statistical significance.