Data Analysis Fundamentals Flashcards
6 cards from real DAC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Data Analysis Fundamentals flashcards as text
What are the four types of data analytics?
Answer: Descriptive (what happened), diagnostic (why), predictive (what will happen), and prescriptive (what should we do)
Analytics progresses from descriptive (summarizing past data), to diagnostic (identifying causes), to predictive (forecasting future outcomes), to prescriptive (recommending actions), each building on the previous level.
What is the difference between structured and unstructured data?
Answer: Structured data is organized in rows/columns (databases); unstructured lacks predefined format (text, images, video)
Structured data fits neatly into relational databases with defined schemas (spreadsheets, SQL databases). Unstructured data lacks a predefined format (emails, social media posts, images, videos) and requires different processing approaches.
What is a p-value in statistical hypothesis testing?
Answer: The probability of obtaining results at least as extreme as observed, assuming the null hypothesis is true
A p-value represents the probability of observing results as extreme as or more extreme than the data, given that the null hypothesis is true. A p-value below a chosen threshold (typically 0.05) leads to rejecting the null hypothesis.
What is ETL in data analytics?
Answer: Extract, Transform, Load — the process of moving data from sources, cleaning/transforming it, and loading into a data warehouse
ETL is the process of extracting data from various sources, transforming it (cleaning, standardizing, aggregating), and loading it into a target data warehouse or analytics platform for analysis.
What is the difference between correlation and causation?
Answer: Correlation measures the statistical relationship between variables; causation means one variable directly causes changes in another
Correlation indicates two variables move together (positive or negative relationship), but causation requires proof that one variable directly influences the other. Confounding variables can create spurious correlations.
What is data cleaning and why is it important?
Answer: Identifying and correcting errors, inconsistencies, and missing values in datasets; critical because analysis quality depends on data quality
Data cleaning (data wrangling) involves handling missing values, correcting errors, standardizing formats, removing duplicates, and resolving inconsistencies — typically consuming 60-80% of an analyst's time but essential for reliable results.