Data Analytics Flashcards
7 cards from real CPA practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analytics flashcards as text
What does ETL stand for in data analytics pipelines?
Answer: Extract, Transform, Load
ETL stands for Extract, Transform, Load — the three-phase process of pulling data from sources, reshaping it, and loading it into a target system.
A data analyst wants to identify which rows in a pandas DataFrame contain null values. Which method should they use?
Answer: df.isnull()
df.isnull() returns a boolean DataFrame of the same shape where True indicates a missing (null) value.
In statistics, what does a p-value less than 0.05 typically indicate?
Answer: The result is statistically significant at the 5% level
A p-value < 0.05 means there is less than a 5% probability of observing the result (or more extreme) if the null hypothesis were true, so we reject it.
Which SQL keyword is used to remove duplicate rows from a query result?
Answer: DISTINCT
The DISTINCT keyword in a SELECT statement eliminates duplicate rows, returning only unique combinations of the selected columns.
What type of data is represented by the Python list ['red', 'blue', 'green', 'red']?
Answer: Nominal categorical data
Colors are nominal categorical data — they represent distinct categories with no inherent order or numerical meaning.
In a box plot, what do the whiskers typically represent?
Answer: The range within 1.5 × IQR from the quartiles
By Tukey's convention, box plot whiskers extend to the farthest data point within 1.5 × IQR beyond Q1 and Q3; points beyond are plotted as outliers.
Which technique splits a dataset into training and testing subsets to evaluate a predictive model's performance?
Answer: Train-test split
A train-test split reserves a portion of data (commonly 20-30%) as an unseen test set to measure how well a trained model generalizes.