โ† All MS-DS Master of Data science Flashcard Decks

Machine Learning Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Machine Learning flashcards as text
  1. What is the key difference between bagging and boosting ensemble methods?

    Answer: Bagging trains base learners in parallel on bootstrap samples; boosting trains them sequentially, each correcting prior errors

    Bagging reduces variance by averaging independently trained models on bootstrap samples, while boosting reduces bias by sequentially fitting models that focus on previously misclassified instances.

  2. In a convolutional neural network (CNN), what does a pooling layer do?

    Answer: Reduces spatial dimensions by aggregating features within local regions

    Pooling layers (max or average) downsample feature maps spatially, reducing computational cost and providing translation invariance by summarizing local regions.

  3. Which evaluation metric is most appropriate for a regression model predicting house prices when outliers are present in the data?

    Answer: Mean Absolute Error (MAE)

    MAE is more robust to outliers than MSE/RMSE because it uses absolute differences rather than squared differences, preventing large errors from disproportionately inflating the metric.

  4. What problem does batch normalization solve in deep network training?

    Answer: Reduces internal covariate shift by normalizing layer inputs to have zero mean and unit variance

    Batch normalization standardizes each layer's inputs across the mini-batch, stabilizing the distribution of activations and allowing higher learning rates and faster convergence.

  5. In the context of Shapley values (SHAP), what does a negative SHAP value for a feature indicate?

    Answer: The feature decreased the model's prediction relative to the average prediction

    A negative SHAP value means that feature's contribution pushed the model's output below the baseline (average prediction), indicating it reduced the predicted value for that instance.

  6. What is the purpose of the attention mechanism in transformer-based models?

    Answer: To compute a weighted combination of all input positions, allowing each position to attend to relevant context globally

    Attention computes query-key similarity scores across all positions, then uses them to weight the values, enabling each token to directly access information from any other token regardless of distance.

  7. Which assumption of linear regression is violated when the residuals' variance increases systematically with the predicted values?

    Answer: Homoscedasticity

    Homoscedasticity requires constant variance of residuals across all levels of predictors; when variance grows with fitted values, this assumption is violated (heteroscedasticity).