Data Science Exam — Questions and Answers
Question 1: You notice that a numeric column has values ranging from 1 to 1,000,000 with most values under 1,000. Which transformation would best reveal structure in the lower range?
- Logarithmic transformation (Correct answer)
- Z-score normalization
- Square root transformation
- Min-max scaling
Correct answer: Logarithmic transformation
A log transformation compresses large values while expanding small values, making structure in the dense lower range visible without losing the high-end variation.
Question 2: A machine learning engineer is using L1 regularization (Lasso) for a linear regression model with a large number of features. What is a key benefit of using L1 regularization in this context?
- It guarantees a more complex and flexible model.
- It can perform automatic feature selection by shrinking some feature coefficients to exactly zero. (Correct answer)
- It is effective at handling multicollinearity by shrinking correlated coefficients together.
- It ensures that all original features are retained in the final model.
Correct answer: It can perform automatic feature selection by shrinking some feature coefficients to exactly zero.
A primary advantage of L1 regularization (Lasso) is its ability to produce sparse models. It adds a penalty equal to the absolute value of the magnitude of coefficients, which can force the coefficients of less important features to become exactly zero, effectively removing them from the model. [2, 11, 28, 29]
Question 3: Under GDPR, the 'right to erasure' (right to be forgotten) means:
- All EU data must be deleted after 5 years automatically
- Companies must erase data after each model training run
- Only government data must be erased on request
- Individuals can demand deletion of their personal data under certain conditions (Correct answer)
Correct answer: Individuals can demand deletion of their personal data under certain conditions
GDPR's right to erasure allows individuals to request deletion of their personal data when it is no longer necessary or consent is withdrawn.
Question 4: Which of the following is the primary advantage of using word embeddings (like Word2Vec) over a one-hot encoding representation?
- They capture semantic relationships, placing words with similar meanings closer in the vector space. (Correct answer)
- They result in sparse, high-dimensional vectors that are easier for models to process.
- They perfectly represent every word in a language without any information loss.
- They are simpler and faster to compute for any given vocabulary.
Correct answer: They capture semantic relationships, placing words with similar meanings closer in the vector space.
Word embeddings represent words as dense, low-dimensional vectors. Their key advantage is that they capture semantic meaning and context. Words with similar meanings (e.g., 'king' and 'queen') will have vectors that are close to each other in the vector space. One-hot encoding creates sparse, high-dimensional vectors that are orthogonal to each other, meaning it cannot capture any notion of similarity between words.
Question 5: What does the power of a hypothesis test measure?
- Probability that the confidence interval contains the true parameter
- Probability of failing to reject Hâ‚€ when Hâ‚€ is false
- Probability of rejecting Hâ‚€ when Hâ‚€ is true
- Probability of rejecting Hâ‚€ when Hâ‚€ is false (Correct answer)
Correct answer: Probability of rejecting Hâ‚€ when Hâ‚€ is false
Power = 1 − β, where β is the Type II error rate; it is the probability of correctly detecting a true effect.
Question 6: In stratified k-fold cross-validation, what property is preserved in each fold?
- The class distribution (Correct answer)
- The model hyperparameters
- The feature variance
- The order of samples
Correct answer: The class distribution
Stratified k-fold ensures each fold reflects the overall class proportions, which is critical for imbalanced datasets.
Question 7: What is 'catastrophic forgetting' in the context of fine-tuning language models?
- The model loses previously learned knowledge when updated on new task data (Correct answer)
- The model fails to converge during pre-training
- The optimizer forgets momentum values between epochs
- The model forgets the fine-tuning task after deployment
Correct answer: The model loses previously learned knowledge when updated on new task data
Catastrophic forgetting occurs when fine-tuning on a new task overwrites the weights learned during pre-training, degrading performance on original capabilities.
Question 8: Which parameter in DBSCAN controls the minimum density required for a region to be considered a core region?
- bandwidth
- epsilon (ε)
- n_clusters
- min_samples (Correct answer)
Correct answer: min_samples
min_samples specifies how many points must be within the epsilon radius for a point to qualify as a core point, directly controlling density thresholds.
Question 9: What does the 'I' in ARIMA stand for, and what does it indicate about the model?
- Independent — that the error terms are independent of each other
- Inversed — that the series has been log-transformed
- Integrated — the number of differencing operations applied to achieve stationarity (Correct answer)
- Interval — the time interval between observations
Correct answer: Integrated — the number of differencing operations applied to achieve stationarity
In ARIMA (AutoRegressive Integrated Moving Average), 'I' stands for Integrated, referring to the order d — how many times the series is differenced to become stationary.
Question 10: When should Mean Absolute Percentage Error (MAPE) be avoided as an evaluation metric?
- When the target variable has large values
- When comparing across different scales
- When true values are near or equal to zero (Correct answer)
- When the dataset contains outliers
Correct answer: When true values are near or equal to zero
MAPE divides by the true value, so near-zero targets cause division by near-zero and produce extremely large or undefined errors.
Question 11: An AI model's training data was collected without proper consent from social media users. Publishing a paper using this model raises which primary ethical concern?
- Lack of hyperparameter tuning
- Violating users' privacy and autonomy through non-consensual data use (Correct answer)
- Overfitting due to social media noise
- Insufficient model regularization
Correct answer: Violating users' privacy and autonomy through non-consensual data use
Using data collected without consent violates individuals' autonomy and privacy rights, regardless of how the data was technically obtained.
Question 12: What is 'data skew' in the context of distributed big data processing?
- Data that fails schema validation
- Data that arrives out of chronological order
- Uneven distribution of data across partitions causing some tasks to run much longer (Correct answer)
- Corrupted records in a dataset
Correct answer: Uneven distribution of data across partitions causing some tasks to run much longer
Data skew occurs when some partitions hold significantly more data than others, creating bottleneck tasks that slow down the entire job.
Question 13: What does ROC-AUC measure in a binary classification model?
- The harmonic mean of precision and recall
- The model's ability to rank positive instances higher than negatives across all thresholds (Correct answer)
- The ratio of true positives to all actual positives
- The exact probability threshold that maximizes accuracy
Correct answer: The model's ability to rank positive instances higher than negatives across all thresholds
AUC represents the probability that a randomly chosen positive example is ranked higher than a randomly chosen negative one.
Question 14: The IEEE Ethically Aligned Design framework prioritizes which of the following as its first principle?
- Efficiency and scalability
- Intellectual property protection
- Open-source accessibility
- Human well-being as the primary goal of AI development (Correct answer)
Correct answer: Human well-being as the primary goal of AI development
IEEE's Ethically Aligned Design places human well-being at the center, asserting that AI must ultimately serve and enhance human flourishing.
Question 15: Which algorithm is a non-parametric, instance-based learning method that classifies new points based on the majority label among their nearest neighbors?
- Support Vector Machine
- Logistic Regression
- Naive Bayes
- K-Nearest Neighbors (KNN) (Correct answer)
Correct answer: K-Nearest Neighbors (KNN)
KNN stores training examples and classifies new instances by a vote among the k closest training points in feature space.
Question 16: A researcher tests 20 independent null hypotheses, each at α = 0.05. The expected number of false rejections under all null hypotheses being true is:
- 0
- 1 (Correct answer)
- 20
- 5
Correct answer: 1
Expected false positives = 20 × 0.05 = 1, illustrating the multiple testing problem even when every null is actually true.
Question 17: How is the Akaike Information Criterion (AIC) used in ARIMA model selection?
- It tests whether the series is stationary before differencing
- It measures out-of-sample forecast accuracy on a held-out test set
- It detects and flags outliers in the training data
- It balances goodness of fit against model complexity to help avoid overfitting (Correct answer)
Correct answer: It balances goodness of fit against model complexity to help avoid overfitting
AIC penalizes adding more parameters while rewarding model fit, guiding selection toward the most parsimonious ARIMA model that adequately captures the data.
Question 18: What language is utilized in the field of data science?
- c++
- Ruby (Correct answer)
- Java
- R
Correct answer: Ruby
While Python and R are dominant in data science, Ruby is a general-purpose language that can also be utilized. It offers capabilities for data manipulation, scripting, and web development, which can be relevant in data-related projects, especially when integrating with web applications. Although its ecosystem for advanced statistical modeling and machine learning is less extensive than Python or R, Ruby's flexibility allows for its use in various data science tasks.
Question 19: In Edward Tufte's framework, what does a high 'data-ink ratio' indicate about a visualization?
- The chart has a dark background
- Most of the ink conveys actual data rather than non-essential embellishments (Correct answer)
- The dataset has many variables
- More decorative elements are used
Correct answer: Most of the ink conveys actual data rather than non-essential embellishments
Tufte's data-ink ratio principle states that unnecessary ink (gridlines, borders, decorations) should be minimized so the visualization focuses attention on the data itself.
Question 20: What does the term 'feature engineering' refer to in data science?
- Creating or transforming input variables to improve model performance (Correct answer)
- Selecting the best algorithm for a task
- Tuning hyperparameters of a model
- Evaluating model accuracy on test data
Correct answer: Creating or transforming input variables to improve model performance
Feature engineering involves creating new features or modifying existing ones to help the model learn better.
Question 21: What is the key advantage of using pre-trained contextual embeddings (like those from BERT) over static word embeddings (like word2vec)?
- They only need character-level input
- They eliminate the need for tokenization
- They produce different representations for the same word in different contexts (Correct answer)
- They require less computational memory
Correct answer: They produce different representations for the same word in different contexts
Contextual embeddings dynamically encode a word's meaning based on its surrounding context, addressing polysemy that static embeddings cannot.
Question 22: Which storage format is columnar and commonly used in the Hadoop ecosystem for analytical workloads?
- CSV
- Parquet (Correct answer)
- Avro
- JSON
Correct answer: Parquet
Parquet is a columnar storage format that provides efficient compression and encoding, making it well-suited for analytical queries on large datasets.
Question 23: Which plot is most suitable for visualizing changes in a continuous variable over time?
- Line chart (Correct answer)
- Box plot
- Histogram
- Pie chart
Correct answer: Line chart
A line chart connects sequential data points in chronological order, making trends, cycles, and anomalies in time-series data easy to see.
Question 24: In a ROC curve for a binary classifier, what two rates are plotted on the axes?
- True Positive Rate vs. False Positive Rate (Correct answer)
- Precision vs. Recall
- Accuracy vs. Loss
- F1-Score vs. Threshold
Correct answer: True Positive Rate vs. False Positive Rate
A ROC curve plots the True Positive Rate (sensitivity) on the Y-axis against the False Positive Rate (1-specificity) on the X-axis across all classification thresholds.
Question 25: "Which machine learning algorithm uses the bagging concept as its foundation?"
- Regression
- Random-forest (Correct answer)
- Classification
- Decision tree
Correct answer: Random-forest
The Random Forest algorithm is an ensemble learning method that fundamentally relies on the bagging (Bootstrap Aggregating) concept. It constructs multiple decision trees during training, each built on a random subset of the training data with replacement. By combining the predictions from these numerous individual trees, Random Forest significantly reduces variance and helps prevent overfitting, leading to more robust and accurate models.
Question 26: Which of the following best describes the primary goal of Explainable AI (XAI)?
- To achieve the highest possible prediction accuracy regardless of model complexity.
- To reduce the computational cost and time required for training complex models.
- To allow stakeholders to understand, trust, and manage the results of an AI model. (Correct answer)
- To fully automate the process of feature selection and data preprocessing.
Correct answer: To allow stakeholders to understand, trust, and manage the results of an AI model.
The core purpose of Explainable AI (XAI) is to make the decision-making process of AI models, especially complex 'black-box' models, transparent and understandable to humans. This fosters trust, ensures accountability, and allows for effective auditing and debugging.
Question 27: What is the primary purpose of the 'gap statistic' method in clustering?
- To detect outliers that fall between clusters
- To compare within-cluster dispersion to a null reference distribution to find optimal k (Correct answer)
- To measure the gap between cluster centroids
- To fill missing values before clustering
Correct answer: To compare within-cluster dispersion to a null reference distribution to find optimal k
The gap statistic compares WCSS of the actual clustering to that expected under a null uniform distribution, with the optimal k being where the gap is maximized.
Question 28: A data science team is analyzing a dataset with many features (high dimensionality). To simplify the data and aid in visualization, they apply a technique that reduces the number of variables while retaining most of the original information. Which EDA technique are they most likely using?
- Data profiling
- Outlier detection
- Hypothesis testing
- Principal Component Analysis (PCA) (Correct answer)
Correct answer: Principal Component Analysis (PCA)
Principal Component Analysis (PCA) is a classic dimensionality reduction technique used in EDA. It transforms a large set of correlated variables into a smaller set of uncorrelated variables called principal components, making it easier to visualize and analyze high-dimensional data without significant information loss.
Question 29: Which probability distribution is most appropriate for modeling the number of events occurring in a fixed time interval?
- Poisson distribution (Correct answer)
- Exponential distribution
- Normal distribution
- Binomial distribution
Correct answer: Poisson distribution
The Poisson distribution models the count of independent events occurring in a fixed interval of time or space.
Question 30: What distinguishes zero-shot classification from few-shot classification in NLP?
- Zero-shot uses larger models than few-shot
- Zero-shot provides no examples; few-shot provides a small number of examples in the prompt (Correct answer)
- Zero-shot fine-tunes the model; few-shot does not
- Zero-shot uses labeled data; few-shot uses unlabeled data
Correct answer: Zero-shot provides no examples; few-shot provides a small number of examples in the prompt
Zero-shot classification relies only on task descriptions or label names, while few-shot provides a handful of labeled examples in the input.
Question 31: What is the role of the kernel in a Support Vector Machine (SVM)?
- It reduces the number of support vectors to speed up inference
- It maps data into a higher-dimensional space to find a linear separating hyperplane (Correct answer)
- It regularizes the margin to prevent overfitting
- It initializes the weights of the model before training
Correct answer: It maps data into a higher-dimensional space to find a linear separating hyperplane
Kernel functions compute dot products in a transformed feature space without explicitly computing the transformation.
Question 32: In supervised learning, what is the 'label'?
- The model's predicted value
- A feature column name
- A data cleaning step
- The known output used during training (Correct answer)
Correct answer: The known output used during training
The label (or target) is the known correct output that the model learns to predict during training.
Question 33: What is the main advantage of using a KDE (Kernel Density Estimate) over a histogram for EDA?
- It is computationally faster to compute
- It requires no assumption about the underlying distribution
- It produces a smooth continuous density curve that is less sensitive to bin width choices (Correct answer)
- It shows exact counts of observations in each interval
Correct answer: It produces a smooth continuous density curve that is less sensitive to bin width choices
KDE creates a smooth density estimate that avoids the arbitrary bin-width problem of histograms, giving a clearer view of the distribution shape.
Question 34: A team has built a binary classification model and visualized its performance using a confusion matrix. The matrix shows the following: True Positives (TP) = 80, False Positives (FP) = 20, True Negatives (TN) = 450, False Negatives (FN) = 50. What is the precision of this model?
- 0.615
- 0.800 (Correct answer)
- 0.900
- 0.937
Correct answer: 0.800
Precision is calculated as the ratio of correctly predicted positive observations to the total predicted positive observations. The formula is TP / (TP + FP). In this case, Precision = 80 / (80 + 20) = 80 / 100 = 0.80. This means that when the model predicts the positive class, it is correct 80% of the time.
Question 35: In coreference resolution, what is the task the model must solve?
- Translating pronouns across languages
- Identifying the sentiment of pronouns
- Detecting grammatical errors in pronoun usage
- Clustering mentions in text that refer to the same real-world entity (Correct answer)
Correct answer: Clustering mentions in text that refer to the same real-world entity
Coreference resolution groups all mentions (nouns, pronouns, noun phrases) that refer to the same entity into clusters.
Question 36: In K-Means, what is the objective function being minimized?
- Maximum distance between any two points in the same cluster
- Sum of pairwise distances between all cluster centroids
- Total between-cluster sum of squares
- Total within-cluster sum of squared distances to centroids (Correct answer)
Correct answer: Total within-cluster sum of squared distances to centroids
K-Means minimizes the Within-Cluster Sum of Squares (WCSS), also called inertia, which is the sum of squared Euclidean distances from each point to its assigned centroid.
Question 37: Select if the following assertion is accurate or not:
- Maybe true of false
- Cannot be determined
- True
- False (Correct answer)
Correct answer: False
Without an assertion provided in the question, it's impossible to determine the specific reason why 'False' is the correct answer. This indicates a flaw in the question itself, as a statement is required to evaluate its accuracy. However, in a typical data science context, 'False' would be chosen if the unstated assertion presented a common misconception or an inaccurate claim about the field.
Question 38: The Cramér-Rao lower bound gives the minimum variance achievable by:
- A sufficient statistic under all loss functions
- The maximum likelihood estimator only
- Any consistent estimator with large n
- Any unbiased estimator of a parameter (Correct answer)
Correct answer: Any unbiased estimator of a parameter
The CRLB states that Var(θ̂) ≥ 1/I(θ) for any unbiased estimator, where I(θ) is the Fisher information.
Question 39: Which NLP task involves identifying the grammatical role of each word in a sentence, such as noun or verb?
- Part-of-speech tagging (Correct answer)
- Semantic role labeling
- Dependency parsing
- Named entity recognition
Correct answer: Part-of-speech tagging
Part-of-speech tagging assigns grammatical categories (noun, verb, adjective, etc.) to each token in a sentence.
Question 40: You have a dataset with 10 features and want to explore relationships between all pairs. How many scatter plots would a full scatter plot matrix contain?
- 90
- 100 (Correct answer)
- 10
- 45
Correct answer: 100
A scatter plot matrix for n variables contains n² cells total, so 10×10 = 100 cells (10 diagonal + 90 off-diagonal scatter plots).
Question 41: A machine learning model exhibits high variance but low bias. What is the most likely characteristic of this model's performance?
- The model is underfitting the training data.
- The model performs well on both the training and test data.
- The model makes strong, simplistic assumptions about the data.
- The model is highly sensitive to small fluctuations in the training data, likely causing overfitting. (Correct answer)
Correct answer: The model is highly sensitive to small fluctuations in the training data, likely causing overfitting.
The bias-variance tradeoff is a central concept in machine learning. A model with high variance is highly flexible and captures a lot of the detail in the training data, including the noise. This sensitivity to the training data means its performance can fluctuate significantly with different training sets, a classic sign of overfitting. Low bias means the model's average prediction is close to the correct value, but the high variance leads to poor generalization on unseen data.
Question 42: In gradient boosting for classification, what does each successive tree learn?
- A linear combination of all previous trees
- An independent model on a bootstrapped dataset
- A separate random subset of features
- The residual errors (pseudo-residuals) of the previous ensemble (Correct answer)
Correct answer: The residual errors (pseudo-residuals) of the previous ensemble
Each new tree in gradient boosting is fit to the negative gradient (pseudo-residuals) of the loss function from the current ensemble.
Question 43: The false discovery rate (FDR), controlled by the Benjamini-Hochberg procedure, is defined as:
- One minus the familywise error rate
- The expected proportion of rejected nulls that are false rejections (Correct answer)
- The probability of at least one false rejection
- The probability that a specific rejected null is true
Correct answer: The expected proportion of rejected nulls that are false rejections
FDR = E[V/R] where V is false rejections and R is total rejections; BH controls this expected proportion, offering more power than Bonferroni.
Question 44: What type of relationship does Pearson correlation measure?
- Rank-based association between ordinal variables
- Causal relationship between variables
- Linear relationship between two continuous variables (Correct answer)
- Any monotonic relationship between two variables
Correct answer: Linear relationship between two continuous variables
Pearson correlation measures the strength and direction of the linear association between two continuous variables, ranging from -1 to +1.
Question 45: What is the role of the key, query, and value matrices in scaled dot-product attention?
- They project inputs to compute attention scores and weighted value outputs (Correct answer)
- They apply layer normalization before softmax
- They store the vocabulary embeddings for lookup
- They encode positional information for each token
Correct answer: They project inputs to compute attention scores and weighted value outputs
Queries and keys are compared via dot product to produce attention weights, which then blend value vectors into the output.
Data Science Exam
The Data Science Examination (DSE) evaluates proficiency in computer science, mathematics, and statistics for data science professionals, administered by Pearson VUE.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds