Mixed Deck — All Data Science Topics Flashcards
100 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 20 Mixed Deck — All Data Science Topics flashcards as text
What is multicollinearity in a regression model?
Answer: High correlation among predictor variables
Multicollinearity occurs when independent variables are highly correlated with each other.
When splitting data into training and test sets, why should stratified sampling be used for imbalanced classification problems?
Answer: To ensure both sets maintain the same class distribution as the original data
Stratified sampling preserves the original class proportions in both splits, preventing situations where the minority class is underrepresented or absent in one set.
Combining two DataFrames on a shared key column in pandas is done with:
Answer: pd.merge()
pd.merge joins DataFrames on common key columns.
When you send a MapReduce task to a Hadoop cluster, you discover that even though it was sent successfully, the job has not yet been finished. What ought you to do?
Answer: Ensure that the TaskTracker is running
In older Hadoop 1.x architectures, the TaskTracker was responsible for executing individual Map and Reduce tasks on the data nodes. If a MapReduce job has been successfully sent but is not finishing, it indicates that the tasks themselves are not being executed. Therefore, ensuring the TaskTracker (or its equivalent, the NodeManager in Hadoop 2.x+) is running is crucial for job execution.
Which dimensionality reduction technique is unsupervised and linear?
Answer: Principal Component Analysis (PCA)
PCA is an unsupervised, linear method that projects data onto directions of maximum variance.
In a neural network, what is the role of the activation function?
Answer: To introduce non-linearity into the model
Activation functions like ReLU or sigmoid introduce non-linearity, allowing neural networks to learn complex, non-linear relationships that a purely linear model could not capture.
An analyst creates a bar chart to compare the monthly revenue of four different products. To better highlight the differences, the y-axis is set to start at $500,000 instead of $0, as all products generated over $520,000. What is the primary issue with this visualization choice?
Answer: It visually exaggerates the proportional differences between the products, potentially misleading the audience.
Truncating the y-axis (i.e., not starting at zero) is a common way to create a misleading graph. [20, 25, 26] For bar charts, the length of the bar is a key pre-attentive attribute that viewers use to compare values. Starting the axis at a non-zero value distorts this visual comparison, making small differences appear much larger and more significant than they actually are. [25, 26]
Which type of data visualization is best suited for showing the distribution of a single continuous variable?
Answer: Histogram
Histograms display the frequency distribution of continuous data by grouping values into bins and showing their counts.
Which of the following best describes the concept of 'feature engineering'?
Answer: Transforming or creating input variables to improve model performance
Feature engineering involves creating new input variables or transforming existing ones — such as encoding categoricals, scaling numerics, or combining columns — to provide the model with more informative signals.
What is the purpose of using stratified sampling when splitting a dataset into training and test sets?
Answer: To maintain the same class distribution in both the training and test sets
Stratified sampling preserves the proportion of each class in both splits, which is critical for imbalanced datasets where random splitting could leave minority classes underrepresented in one set.
What role does the Reduce function play in the MapReduce framework?
Answer: It distributes the input to multiple nodes for processing.
In the MapReduce framework, the Reduce function processes the intermediate key-value pairs generated by the Map tasks. Before reaching the Reduce function, these intermediate results are shuffled and sorted, effectively distributing them to the appropriate reducer tasks across multiple nodes. Thus, the Reduce function is a crucial component in the overall distributed processing pipeline, working on data that has been distributed for aggregation.
A data scientist is preparing a dataset with a categorical feature "City" ('New York', 'London', 'Tokyo'). The feature has high cardinality and no inherent order. Which encoding technique is most appropriate to convert this feature for a linear model without imposing a false ordinal relationship?
Answer: One-Hot Encoding
One-Hot Encoding is the correct method for nominal categorical variables (where no order exists) that will be used in linear models. It creates new binary columns for each category, representing its presence or absence without implying any sort of ranking, which would be an incorrect assumption for a feature like 'City'.
What is a key limitation of using autoencoders for dimensionality reduction compared to PCA?
Answer: Autoencoders can overfit and require careful regularization and tuning
Autoencoders are neural networks that can memorize training data if not properly regularized, making them prone to overfitting on small datasets.
A polynomial feature of degree 2 on feature x adds which term?
Answer: x squared
Degree-2 polynomial expansion adds x² (and cross terms for multiple features).
What type of plot uses rectangular bars whose heights represent the frequency of data falling within specified ranges?
Answer: Histogram
A histogram displays the distribution of a single continuous variable by dividing data into intervals (bins) and showing the count or frequency of observations in each bin.
Which chart type is most appropriate for showing the part-to-whole composition of a single category at one point in time?
Answer: Pie chart
A pie chart shows how parts make up a whole for a single categorical breakdown.
What is the purpose of batch normalization in a neural network?
Answer: To normalize layer inputs to reduce internal covariate shift and speed up training
Batch normalization normalizes activations within each mini-batch, stabilizing training and often allowing higher learning rates.
When would a Wilcoxson Rank Sum test be appropriate?
Answer: When you cannot make an assumption about the distribution of the populations
The Wilcoxon Rank Sum test, also known as the Mann-Whitney U test, is a non-parametric statistical test. It is appropriate when comparing two independent groups and you cannot make assumptions about the underlying distribution of the populations, such as normality. This test ranks the data and compares the sums of the ranks between the groups.
Which optimization algorithm adapts the learning rate for each parameter using estimates of first and second moments of the gradients?
Answer: Adam
Adam (Adaptive Moment Estimation) combines momentum and RMSProp by tracking both the first and second moments of gradients.
Which is a key reason to choose a horizontal bar chart over vertical bars?
Answer: Long category labels are easier to read horizontally
Horizontal bars give room for long category labels without rotating or truncating text.