FREE Data Science Analysis Question and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 FREE Data Science Analysis Question and Answers flashcards as text
Which cross-validation strategy is most appropriate for time series data?
Answer: Time-based forward chaining (walk-forward validation)
Forward chaining respects the temporal order of observations by always training on past data and validating on future data, preventing data leakage.
What does a high variance inflation factor (VIF) for a predictor variable indicate in multiple regression?
Answer: The predictor is highly correlated with other predictors in the model
A high VIF indicates multicollinearity, meaning the predictor can be largely explained by a linear combination of other predictors in the model.
When building a data pipeline, what is the primary advantage of using a medallion architecture (bronze, silver, gold layers)?
Answer: It separates raw ingestion from cleaned and aggregated data, improving data quality and traceability
The medallion architecture progressively refines data through distinct layers, making it easier to trace lineage, debug issues, and serve analytics-ready datasets.
In hypothesis testing, what does a Type II error represent?
Answer: Failing to reject a false null hypothesis
A Type II error occurs when the test fails to detect a real effect, meaning the null hypothesis is not rejected even though it is actually false.
Which technique is best suited for identifying anomalies in a high-dimensional dataset where normal data forms irregular clusters?
Answer: Isolation Forest
Isolation Forest efficiently isolates anomalies by randomly partitioning features, and it works well in high-dimensional spaces without assuming a specific data distribution.
What is the purpose of applying a log transformation to a heavily right-skewed feature before modeling?
Answer: To compress the range of extreme values and make the distribution more symmetric
Log transformation reduces the influence of extreme high values and pulls the right tail inward, producing a more symmetric distribution that better satisfies model assumptions.