Mixed Deck — All MS-DS Master of Data science Topics Flashcards
100 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 20 Mixed Deck — All MS-DS Master of Data science Topics flashcards as text
What is the primary purpose of the kernel trick in Support Vector Machines?
Answer: To implicitly map input features into a higher-dimensional space without computing the transformation explicitly
The kernel trick allows SVMs to find a linear separating hyperplane in a higher-dimensional feature space by computing inner products via a kernel function, avoiding the computational cost of explicit transformation.
What is the purpose of A/B testing in a data-driven organization?
Answer: To compare two versions of a variable to determine which performs better
A/B testing is a controlled experiment where two variants are compared to measure which one produces a better outcome on a specific metric.
When comparing more than two group means simultaneously, a one-way ANOVA is preferred over multiple t-tests because:
Answer: Multiple t-tests inflate the Type I error rate
Running multiple t-tests increases the familywise error rate; ANOVA controls it by testing all groups in a single analysis.
In the context of Shapley values (SHAP), what does a negative SHAP value for a feature indicate?
Answer: The feature decreased the model's prediction relative to the average prediction
A negative SHAP value means that feature's contribution pushed the model's output below the baseline (average prediction), indicating it reduced the predicted value for that instance.
Which test is most appropriate to assess whether two continuous variables are linearly associated when both are normally distributed?
Answer: Pearson correlation test
Pearson correlation tests linear association between two normally distributed continuous variables and is parametric.
A researcher uses bootstrapping to estimate the standard error of the median. What does this method fundamentally rely on?
Answer: Resampling with replacement from the observed data to approximate the sampling distribution
Bootstrapping treats the empirical distribution as a proxy for the population and resamples with replacement to build an approximate sampling distribution.
Which of the following suggests there is no association between any of the relationships?
Answer: Cor(X, Y) = 0
The correlation coefficient, Cor(X, Y), measures the strength and direction of the linear relationship between two random variables X and Y. A value of 0 indicates that there is no linear association between the variables. It's important to note that a correlation of 0 does not necessarily imply independence, as non-linear relationships might still exist.
Which regularization technique specific to transformer training helps stabilize optimization by normalizing activations before each sublayer's computation?
Answer: Pre-layer normalization (Pre-LN) applied before each sublayer
Pre-LN applies layer normalization to the input of each sublayer (attention or FFN) before the residual addition, leading to more stable gradient flow compared to the original Post-LN Transformer.
The problems with big data veracity go beyond volume, diversity, and velocity.
Answer: True
The '3 V's' (Volume, Velocity, Variety) are commonly associated with Big Data, but 'Veracity' is a fourth crucial dimension. Veracity refers to the trustworthiness, accuracy, and quality of the data, addressing issues like bias, noise, and abnormalities that can significantly impact analysis and decision-making.
In maximum likelihood estimation, what is the Fisher information used to approximate?
Answer: The variance of the MLE asymptotically
The inverse of the Fisher information provides the asymptotic variance of the MLE, forming the basis of the Cramér-Rao lower bound.
In Bayesian inference, the posterior distribution is proportional to:
Answer: The likelihood times the prior
By Bayes' theorem, P(θ|data) ∝ P(data|θ) × P(θ), i.e., posterior ∝ likelihood × prior.
A data scientist trains a decision tree model. They observe that the model achieves 99% accuracy on the training data but only 70% accuracy on the test data. The performance on the training data is excellent, but the performance on the test data is significantly worse. This situation is a classic example of what?
Answer: High variance (Overfitting)
High variance, or overfitting, occurs when a model learns the training data too well, including its noise and random fluctuations, rather than the underlying general patterns. This results in excellent performance on the training data but poor performance on new, unseen data (the test set).
The degrees of freedom in the Chi-squared distribution are twice as many.
Answer: variance
For a Chi-squared distribution, the degrees of freedom (often denoted as 'k') define its shape and properties. The mean of a Chi-squared distribution is equal to its degrees of freedom (k), and its variance is equal to twice its degrees of freedom (2k). Therefore, the degrees of freedom are directly related to, and half of, the variance.
In a random forest, what is the purpose of feature randomness (selecting a subset of features at each split)?
Answer: It decorrelates the individual trees, reducing ensemble variance
By considering only a random subset of features at each split, trees in a random forest are de-correlated, so their errors are less likely to coincide, reducing overall variance.
Which unsupervised learning technique reduces dimensionality by finding orthogonal axes that maximize variance in the data?
Answer: Principal Component Analysis (PCA)
PCA identifies principal components as orthogonal directions of maximum variance, enabling dimensionality reduction while preserving the most information.
When constructing a likelihood ratio test, what are you comparing?
Answer: The likelihood under the null hypothesis to the likelihood under the alternative
A likelihood ratio test compares the maximized likelihood under the restricted null model to the maximized likelihood under the full alternative model.
Which cross-validation strategy is most appropriate when working with time-series data in a predictive modeling research project?
Answer: Time-series split (expanding window) cross-validation
Time-series split preserves the temporal ordering of observations, ensuring the model is always trained on past data and tested on future data.
What is the primary advantage of using the Matthews Correlation Coefficient (MCC) over accuracy for binary classification on imbalanced datasets?
Answer: MCC accounts for all four confusion matrix categories proportionally
MCC uses all four confusion matrix values (TP, TN, FP, FN) and produces a balanced measure even when classes are highly imbalanced.
Which storytelling framework is commonly used to structure a data presentation narrative?
Answer: Situation-Complication-Resolution
The SCR framework guides presenters to establish context, highlight the problem, and propose a data-driven solution.
In the context of association rule mining, what does the 'lift' metric indicate?
Answer: The ratio of observed support to expected support if items were independent
Lift measures how much more often items appear together than expected if they were statistically independent, with values greater than 1 indicating a positive association.