Basic Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Basic flashcards as text
What is the difference between structured and unstructured data?
Answer: Structured data fits in predefined schemas; unstructured does not
Structured data is organized in predefined schemas (like tables), while unstructured data lacks a fixed format (like images or emails).
Which sampling method gives every member of the population an equal chance of being selected?
Answer: Simple random sampling
Simple random sampling ensures every individual has an equal probability of selection, minimizing selection bias.
What is a confusion matrix used for?
Answer: Evaluating classification model performance
A confusion matrix shows counts of true positives, false positives, true negatives, and false negatives for a classifier.
Which of the following algorithms is used for dimensionality reduction?
Answer: Principal Component Analysis (PCA)
PCA projects high-dimensional data onto fewer principal components that capture the most variance.
In data science, what is a 'data lake'?
Answer: A centralized repository storing raw data in any format at scale
A data lake stores raw, unprocessed data in any format at scale, unlike a data warehouse which stores structured, processed data.
What does a p-value less than 0.05 typically indicate?
Answer: The result is statistically significant at the 5% level
A p-value < 0.05 means there is less than a 5% chance of observing results this extreme if the null hypothesis were true.
Which of the following best describes 'data wrangling'?
Answer: Cleaning and transforming raw data into a usable format
Data wrangling (also called data munging) involves cleaning, restructuring, and enriching raw data to make it analysis-ready.