CompTIA Data+ (DA0-001) — Questions and Answers
Question 1: What does a chi-square test assess?
- The difference between two means
- The association between two categorical variables (Correct answer)
- The correlation between two continuous variables
- The normality of a dataset
Correct answer: The association between two categorical variables
A chi-square test of independence evaluates whether there is a statistically significant association between two categorical variables.
Question 2: Which type of database is best suited for storing unstructured or semi-structured data?
- Relational (RDBMS)
- Hierarchical
- Network
- NoSQL (Correct answer)
Correct answer: NoSQL
NoSQL databases like MongoDB or Cassandra are designed to handle unstructured, semi-structured, and rapidly changing data models.
Question 3: Which aggregate function returns the total number of rows in a result set?
- COUNT() (Correct answer)
- AVG()
- SUM()
- MAX()
Correct answer: COUNT()
COUNT() returns the number of rows that match the specified condition, or total rows when used with *.
Question 4: Which color scale is recommended for sequential data with a meaningful zero point?
- Qualitative palette
- Rainbow color scale
- Sequential single-hue scale (Correct answer)
- Diverging color scale
Correct answer: Sequential single-hue scale
A sequential single-hue scale uses varying lightness of one color to represent ordered data, making magnitude differences clear.
Question 5: What is the purpose of a choropleth map?
- Animating data changes over time
- Using color intensity to represent data values across geographic regions (Correct answer)
- Showing flight paths between cities
- Displaying 3D topographic data
Correct answer: Using color intensity to represent data values across geographic regions
A choropleth map shades geographic regions with colors proportional to a data variable, enabling geographic comparisons.
Question 6: What is a data quality rule?
- A company policy about data ownership
- A SQL constraint on a database column
- A regulatory requirement from a government agency
- A condition or constraint that data must satisfy to be considered fit for use (Correct answer)
Correct answer: A condition or constraint that data must satisfy to be considered fit for use
A data quality rule defines the criteria that data must meet to be considered accurate, complete, and valid for its intended business use.
Question 7: What is the main advantage of Apache Spark over Hadoop MapReduce?
- Spark processes data in-memory, making it significantly faster (Correct answer)
- Spark requires fewer servers
- Spark is easier to install
- Spark uses less disk space
Correct answer: Spark processes data in-memory, making it significantly faster
Apache Spark's in-memory processing can be up to 100x faster than Hadoop MapReduce for iterative algorithms and interactive queries.
Question 8: What is the purpose of data partitioning in big data systems?
- To divide large datasets into smaller chunks for faster parallel processing (Correct answer)
- To replicate data across regions
- To compress data for storage efficiency
- To encrypt sensitive data
Correct answer: To divide large datasets into smaller chunks for faster parallel processing
Partitioning divides data into smaller, manageable chunks that can be processed in parallel across multiple nodes, improving query performance.
Question 9: What is a data mart?
- A real-time streaming database
- A subset of a data warehouse focused on a specific business function or department (Correct answer)
- A marketplace for buying and selling data
- A type of OLTP database
Correct answer: A subset of a data warehouse focused on a specific business function or department
A data mart is a focused subset of a data warehouse tailored to meet the specific reporting and analysis needs of a business department like marketing or finance.
Question 10: What is a stored procedure in SQL?
- A precompiled set of SQL statements stored in the database (Correct answer)
- A backup copy of a table
- An automated query scheduler
- A type of database constraint
Correct answer: A precompiled set of SQL statements stored in the database
A stored procedure is a precompiled, named block of SQL code stored in the database that can be executed on demand.
Question 11: In Power BI, what is a 'slicer' used for?
- Dividing reports into pages
- Slicing columnar data in a table
- Cutting data into training and test sets
- Filtering dashboard visuals interactively by selected values (Correct answer)
Correct answer: Filtering dashboard visuals interactively by selected values
A slicer in Power BI is an interactive filter control that lets users segment all visuals on a report page by selected values.
Question 12: What is the difference between OLTP and OLAP systems?
- OLTP handles real-time transactions; OLAP handles analytical queries on historical data (Correct answer)
- OLTP handles reporting; OLAP handles transactions
- OLTP uses NoSQL; OLAP uses relational databases only
- OLTP is cloud-based; OLAP is on-premise
Correct answer: OLTP handles real-time transactions; OLAP handles analytical queries on historical data
OLTP (Online Transaction Processing) systems handle high-volume day-to-day transactions, while OLAP systems are optimized for complex analytical queries on historical data.
Question 13: What is a fact table in a data warehouse?
- A central table storing measurable, numeric business metrics and foreign keys to dimensions (Correct answer)
- A table containing only textual descriptions
- A lookup table for product codes
- A table containing database configuration settings
Correct answer: A central table storing measurable, numeric business metrics and foreign keys to dimensions
A fact table stores quantitative business metrics (like sales revenue or order quantity) along with foreign keys linking to dimension tables.
Question 14: What is data deduplication?
- Creating redundant database copies for failover
- Encrypting duplicate copies of data
- Backing up data in multiple locations
- The process of identifying and removing duplicate records from datasets (Correct answer)
Correct answer: The process of identifying and removing duplicate records from datasets
Data deduplication identifies and eliminates duplicate records in datasets to ensure each entity appears only once, improving data quality.
Question 15: What does 'data mesh' refer to in modern data architecture?
- A type of mesh network for IoT sensors
- A network topology for data centers
- A decentralized approach where domain teams own and manage their own data as products (Correct answer)
- A centralized data lake with mesh indexing
Correct answer: A decentralized approach where domain teams own and manage their own data as products
Data mesh is a decentralized architectural approach that treats data as a product, with domain-oriented teams owning their own data pipelines.
Question 16: What is ad hoc reporting in business intelligence?
- Reports generated by AI without user input
- Standard reports built by IT departments
- Custom, on-demand reports created by users to answer specific questions (Correct answer)
- Scheduled reports generated automatically at fixed intervals
Correct answer: Custom, on-demand reports created by users to answer specific questions
Ad hoc reporting allows business users to create custom, one-time reports on demand without relying on IT, using self-service BI tools.
Question 17: What does 'drill down' mean in OLAP analysis?
- Removing data from a report
- Moving from a summary view to a more detailed level of data (Correct answer)
- Combining multiple data sources
- Filtering data by a specific dimension
Correct answer: Moving from a summary view to a more detailed level of data
Drilling down navigates from a higher-level summary (e.g., annual sales) to a more granular level (e.g., monthly or daily sales).
Question 18: Which type of chart is best for showing data trends over time?
- Pie chart
- Bar chart
- Line chart (Correct answer)
- Scatter plot
Correct answer: Line chart
A line chart connects data points chronologically, making it ideal for visualizing trends, patterns, and changes over time.
Question 19: What does a box plot (box-and-whisker plot) display?
- Correlation between two variables
- Distribution summary including median, quartiles, and outliers (Correct answer)
- Mean and standard deviation only
- Frequency of each category
Correct answer: Distribution summary including median, quartiles, and outliers
A box plot shows the five-number summary (minimum, Q1, median, Q3, maximum) and visually highlights outliers.
Question 20: Which algorithm builds an ensemble of decision trees using random subsets of features and training samples?
- K-Nearest Neighbors
- Random Forest (Correct answer)
- Naive Bayes
- Gradient Boosting
Correct answer: Random Forest
Random Forest creates multiple decision trees using random feature subsets and bootstrap samples, then aggregates their predictions to improve accuracy and reduce overfitting.
Question 21: What is a Gantt chart primarily used for in data projects?
- Visualizing database schemas
- Displaying statistical correlations
- Project scheduling and timeline management (Correct answer)
- Mapping data lineage
Correct answer: Project scheduling and timeline management
A Gantt chart displays project tasks as horizontal bars along a timeline, showing start dates, durations, and dependencies.
Question 22: Which principle of data visualization does Edward Tufte's 'data-ink ratio' emphasize?
- Maximizing the proportion of ink used to display data (Correct answer)
- Including as many chart types as possible
- Adding gridlines for readability
- Using bright colors for emphasis
Correct answer: Maximizing the proportion of ink used to display data
Tufte's data-ink ratio advocates removing all non-essential chart elements to maximize the proportion of ink conveying actual data.
Question 23: Which visualization tool is most commonly used for business intelligence dashboards in the US?
- Matplotlib
- Tableau (Correct answer)
- Excel pivot tables
- R ggplot2
Correct answer: Tableau
Tableau is the leading BI dashboard tool in the US enterprise market, known for its drag-and-drop interface and interactive dashboards.
Question 24: What is a heat map used to represent?
- Network connections between nodes
- Geographic temperature data only
- Categorical comparisons over time
- Magnitude of values across two dimensions using color intensity (Correct answer)
Correct answer: Magnitude of values across two dimensions using color intensity
A heat map encodes data values as color intensity across a matrix, making it easy to spot patterns and outliers.
Question 25: What is a treemap visualization used for?
- Showing decision tree algorithms
- Tracking project timelines
- Mapping geographic data
- Displaying hierarchical data as nested rectangles sized by value (Correct answer)
Correct answer: Displaying hierarchical data as nested rectangles sized by value
A treemap displays hierarchical data as nested rectangles, where rectangle size represents a quantitative variable.
Question 26: What is a confusion matrix used to measure?
- The convergence rate of the training algorithm
- The variance explained by each principal component
- The correlation between input features
- The counts of true positives, true negatives, false positives, and false negatives (Correct answer)
Correct answer: The counts of true positives, true negatives, false positives, and false negatives
A confusion matrix displays the counts of correct and incorrect predictions broken down by class, enabling calculation of precision, recall, and other classification metrics.
Question 27: Which of the following best describes hyperparameter tuning?
- Adjusting model weights during backpropagation
- Searching for the best configuration settings that control the learning process (Correct answer)
- Increasing the size of the training data
- Removing irrelevant features from the dataset
Correct answer: Searching for the best configuration settings that control the learning process
Hyperparameter tuning involves searching for optimal settings (like learning rate, tree depth, or regularization strength) that are set before training and control how the model learns.
Question 28: Which SQL clause is used to filter rows after aggregation?
- ORDER BY
- WHERE
- GROUP BY
- HAVING (Correct answer)
Correct answer: HAVING
HAVING filters groups created by GROUP BY, whereas WHERE filters individual rows before aggregation.
Question 29: What is the difference between UNION and UNION ALL in SQL?
- UNION allows NULL; UNION ALL does not
- UNION removes duplicates; UNION ALL keeps all rows (Correct answer)
- UNION sorts results; UNION ALL does not
- UNION requires same column count; UNION ALL does not
Correct answer: UNION removes duplicates; UNION ALL keeps all rows
UNION removes duplicate rows from the combined result, while UNION ALL retains all rows including duplicates.
Question 30: What is the primary advantage of using ensemble methods like bagging?
- They combine multiple models to reduce variance and improve predictive performance (Correct answer)
- They eliminate the need for feature engineering
- They reduce the number of hyperparameters to tune
- They train faster than single models on large datasets
Correct answer: They combine multiple models to reduce variance and improve predictive performance
Bagging (Bootstrap Aggregating) trains multiple models on different bootstrap samples and averages their predictions, reducing variance and improving overall model stability.
Question 31: Which of the following is an example of an unsupervised learning algorithm?
- K-means clustering (Correct answer)
- Support vector machine
- Random forest
- Linear regression
Correct answer: K-means clustering
K-means clustering is unsupervised because it groups data points based on similarity without using labeled training examples.
Question 32: What is Apache Hadoop primarily used for?
- Relational database management
- Distributed storage and batch processing of large datasets (Correct answer)
- In-memory data analytics
- Real-time stream processing
Correct answer: Distributed storage and batch processing of large datasets
Apache Hadoop is a distributed computing framework that uses HDFS for storage and MapReduce for batch processing of large datasets across clusters.
Question 33: Which distribution is characterized by a long right tail and is often seen in income data?
- Normal distribution
- Right-skewed (positive skew) distribution (Correct answer)
- Uniform distribution
- Bimodal distribution
Correct answer: Right-skewed (positive skew) distribution
A right-skewed distribution has a tail extending to higher values, common in income data where most people earn moderate amounts but a few earn very high incomes.
Question 34: What does a waterfall chart primarily illustrate?
- Statistical distributions
- Cumulative effect of sequential positive and negative values (Correct answer)
- Hierarchical data structures
- Time series trends
Correct answer: Cumulative effect of sequential positive and negative values
A waterfall chart shows how an initial value is affected by a series of positive and negative incremental changes to reach a final value.
Question 35: What is the purpose of cross-validation in machine learning?
- To visualize model performance
- To increase model training speed
- To remove outliers from training data
- To assess how well a model generalizes to independent datasets (Correct answer)
Correct answer: To assess how well a model generalizes to independent datasets
Cross-validation evaluates model generalization by training and testing on different subsets of data, reducing overfitting risk.
Question 36: What is transfer learning in the context of deep learning?
- Converting a model from one framework to another
- Sending a model from one server to another for inference
- Reusing a pre-trained model on a new but related task (Correct answer)
- Transferring data between databases for training
Correct answer: Reusing a pre-trained model on a new but related task
Transfer learning leverages knowledge from a model trained on a large dataset and fine-tunes it for a new, related task, reducing training time and data requirements.
Question 37: What is the purpose of a data catalog in modern BI?
- A system for archiving old reports
- A catalog of BI report templates
- A searchable inventory of data assets with metadata, definitions, and lineage information (Correct answer)
- A physical storage system for data backups
Correct answer: A searchable inventory of data assets with metadata, definitions, and lineage information
A data catalog helps users discover, understand, and trust data assets by providing metadata, business definitions, data lineage, and usage information.
Question 38: What is a slowly changing dimension (SCD) in data warehousing?
- A real-time dimension updated every second
- A technique for managing historical changes to dimension attributes over time (Correct answer)
- A dimension whose values change very slowly due to database performance issues
- A dimension table with few records
Correct answer: A technique for managing historical changes to dimension attributes over time
SCD (Slowly Changing Dimension) is a method for handling changes to dimension attributes over time, preserving history of how values changed.
Question 39: What is a data breach in the context of compliance?
- An unauthorized access or exposure of protected, sensitive, or confidential data (Correct answer)
- A performance issue causing data loss
- A violation of data retention policy
- A failed database migration
Correct answer: An unauthorized access or exposure of protected, sensitive, or confidential data
A data breach is a security incident where unauthorized individuals access, steal, or expose confidential or protected data.
Question 40: What does GDPR stand for and how does it affect US companies?
- Government Data Protection Registry — applies to federal agencies only
- General Data Protection Regulation — applies to US companies handling EU residents' data (Correct answer)
- General Data Privacy Rule — applies only in Europe
- Global Data Privacy Requirements — applies to all multinational companies
Correct answer: General Data Protection Regulation — applies to US companies handling EU residents' data
GDPR is the EU's data protection regulation that applies to any organization, including US companies, that processes personal data of EU residents.
Question 41: What is data tokenization?
- Replacing sensitive data with non-sensitive placeholder tokens while preserving format (Correct answer)
- Converting data into blockchain tokens
- Encoding data in base64 format
- Splitting datasets into training tokens for ML
Correct answer: Replacing sensitive data with non-sensitive placeholder tokens while preserving format
Tokenization substitutes sensitive data (like credit card numbers) with non-sensitive tokens that have no exploitable value if intercepted.
Question 42: What is predictive analytics?
- Using historical data and statistical models to forecast future outcomes (Correct answer)
- Describing current business conditions
- Analyzing what happened in the past
- Creating real-time dashboards
Correct answer: Using historical data and statistical models to forecast future outcomes
Predictive analytics uses historical data, statistical algorithms, and machine learning to identify the likelihood of future outcomes.
Question 43: What is the California Consumer Privacy Act (CCPA)?
- An FTC regulation on digital advertising
- A national cybersecurity standard for data protection
- A California state law giving consumers rights to know, delete, and opt out of the sale of their personal data (Correct answer)
- A federal law requiring encryption of all consumer data
Correct answer: A California state law giving consumers rights to know, delete, and opt out of the sale of their personal data
CCPA is a California privacy law granting consumers rights to know what personal data is collected, request deletion, and opt out of data sales.
Question 44: What is the purpose of an index in a relational database?
- To speed up data retrieval queries (Correct answer)
- To normalize table structure
- To encrypt stored data
- To enforce referential integrity
Correct answer: To speed up data retrieval queries
An index creates a data structure that allows the database engine to find rows faster without scanning the entire table.
Question 45: What does a scatter plot primarily show?
- Relationship between two numeric variables (Correct answer)
- Proportions of a whole
- Trends over time
- Frequency of categories
Correct answer: Relationship between two numeric variables
A scatter plot displays the correlation or relationship between two continuous numeric variables using dots.
Question 46: What is role-based access control (RBAC) in data systems?
- An encryption method for sensitive data
- A method for assigning random access permissions
- A database replication strategy
- A security model granting data access based on a user's organizational role (Correct answer)
Correct answer: A security model granting data access based on a user's organizational role
RBAC assigns permissions to roles rather than individuals, granting users access to data based on their job function within the organization.
Question 47: What does 'data at rest' refer to in security?
- Data that has not been accessed recently
- Data temporarily cached in memory
- Data waiting to be processed in a queue
- Data stored in databases, file systems, or storage media that is not actively moving (Correct answer)
Correct answer: Data stored in databases, file systems, or storage media that is not actively moving
Data at rest refers to inactive data stored persistently in databases, file systems, or storage media, as opposed to data in transit or in use.
Question 48: What does an ROC curve illustrate in model evaluation?
- The tradeoff between true positive rate and false positive rate at various thresholds (Correct answer)
- The relationship between training loss and validation loss
- The distribution of residuals in a regression model
- The importance of each feature in a classification model
Correct answer: The tradeoff between true positive rate and false positive rate at various thresholds
An ROC curve plots the true positive rate against the false positive rate across different classification thresholds, showing a model's discriminative ability.
Question 49: What is prescriptive analytics?
- Analyzing past performance data
- Describing the current state of business
- Recommending specific actions to achieve desired outcomes based on data models (Correct answer)
- Predicting what will happen next quarter
Correct answer: Recommending specific actions to achieve desired outcomes based on data models
Prescriptive analytics uses optimization and simulation algorithms to recommend specific actions organizations should take to achieve desired outcomes.
Question 50: In a neural network, what is the role of an activation function?
- To normalize input data before training
- To introduce non-linearity so the network can learn complex patterns (Correct answer)
- To initialize the weights of the network
- To reduce the number of parameters in the model
Correct answer: To introduce non-linearity so the network can learn complex patterns
Activation functions introduce non-linearity into neural networks, enabling them to learn and represent complex, non-linear relationships in data.
Question 51: What is a KPI in the context of data visualization?
- Key Performance Indicator (Correct answer)
- Key Process Integration
- Kernel Processing Index
- Known Pattern Insight
Correct answer: Key Performance Indicator
A KPI (Key Performance Indicator) is a measurable value that demonstrates how effectively a company is achieving key business objectives.
Question 52: What is Bayes' Theorem used for?
- Calculating the median of a dataset
- Measuring data spread around the mean
- Testing differences between two populations
- Updating probability estimates based on new evidence (Correct answer)
Correct answer: Updating probability estimates based on new evidence
Bayes' Theorem calculates the conditional probability of an event by updating prior beliefs with new evidence.
Question 53: What is a data dictionary?
- A programming library for data manipulation
- A database index for faster lookups
- A centralized repository documenting data elements, definitions, formats, and relationships (Correct answer)
- A glossary of data science terms for new employees
Correct answer: A centralized repository documenting data elements, definitions, formats, and relationships
A data dictionary documents metadata about data elements including field names, data types, definitions, acceptable values, and relationships.
Question 54: What is a master data management (MDM) system?
- An ETL tool for moving data between systems
- A technology for creating a single, authoritative source of truth for critical business data (Correct answer)
- A database backup and recovery system
- A system for managing database administrator accounts
Correct answer: A technology for creating a single, authoritative source of truth for critical business data
MDM creates and maintains a single, consistent, and accurate record of critical business entities like customers, products, and suppliers across the organization.
Question 55: Which approach is used to handle missing values in a dataset by replacing them with the mean, median, or mode?
- Feature scaling
- One-hot encoding
- Imputation (Correct answer)
- Dimensionality reduction
Correct answer: Imputation
Imputation fills in missing values with a statistical estimate (mean, median, or mode) or a model-based prediction to preserve data completeness for training.
Question 56: What is the role of a data steward in a BI organization?
- To manage cloud infrastructure
- To build ETL pipelines
- To write SQL queries for analysts
- To manage data quality, definitions, and governance policies for assigned data domains (Correct answer)
Correct answer: To manage data quality, definitions, and governance policies for assigned data domains
A data steward oversees data quality, maintains business definitions, and enforces data governance policies for specific data domains within the organization.
Question 57: What does a z-score of 2.0 indicate?
- The value is 2 standard deviations above the mean (Correct answer)
- The value is 2 units below the mean
- The value is twice the median
- The value has a 2% probability of occurring
Correct answer: The value is 2 standard deviations above the mean
A z-score of 2.0 means the data point is 2 standard deviations above the mean of the distribution.
Question 58: What is Google BigQuery's primary architectural advantage?
- On-premise storage for security
- Serverless, scalable analytics separating storage from compute (Correct answer)
- Built-in machine learning only
- Native integration with Excel
Correct answer: Serverless, scalable analytics separating storage from compute
BigQuery's serverless architecture separates storage from compute, allowing independent scaling and eliminating infrastructure management.
Question 59: What is a star schema in data warehousing?
- A database schema designed for transactional processing
- A schema using only a single table
- A central fact table connected to multiple dimension tables (Correct answer)
- A schema with multiple fact tables connected to each other
Correct answer: A central fact table connected to multiple dimension tables
A star schema has a central fact table containing measurable metrics connected to multiple denormalized dimension tables, optimized for analytical queries.
Question 60: What is a view in SQL?
- A virtual table defined by a stored query (Correct answer)
- A backup of the database
- A physical copy of a table
- A type of database index
Correct answer: A virtual table defined by a stored query
A view is a virtual table whose contents are defined by a SELECT query stored in the database.
Question 61: Which SQL command permanently removes a table and its data from the database?
- DELETE
- REMOVE
- TRUNCATE
- DROP (Correct answer)
Correct answer: DROP
DROP TABLE removes the entire table definition and all its data permanently from the database.
Question 62: In logistic regression, what does the output of the sigmoid function represent?
- The error rate of the model
- The distance from the decision boundary
- The exact class label for a data point
- The probability that a data point belongs to a particular class (Correct answer)
Correct answer: The probability that a data point belongs to a particular class
The sigmoid function maps any real value to a probability between 0 and 1, representing the likelihood that an input belongs to the positive class.
Question 63: What does 'self-service BI' mean?
- BI systems that manage their own infrastructure
- Empowering business users to access, analyze, and visualize data without IT assistance (Correct answer)
- Automated reporting without any human involvement
- BI tools that don't require vendor support
Correct answer: Empowering business users to access, analyze, and visualize data without IT assistance
Self-service BI enables non-technical business users to create reports, dashboards, and analyses without depending on IT or data engineers.
Question 64: What is a confidence interval?
- A range of values likely to contain the true population parameter (Correct answer)
- The range within which the sample mean falls
- The margin of error in data collection
- The probability that a hypothesis is correct
Correct answer: A range of values likely to contain the true population parameter
A confidence interval provides a range of plausible values for a population parameter with a specified level of confidence (e.g., 95%).
Question 65: Which metric is most appropriate when evaluating a classification model on a highly imbalanced dataset?
- R-squared
- F1-score (Correct answer)
- Mean squared error
- Accuracy
Correct answer: F1-score
F1-score balances precision and recall, making it a better metric than accuracy for imbalanced datasets where the majority class can dominate the accuracy score.
Question 66: What is Apache Kafka primarily used for?
- Batch processing historical data
- Relational data storage
- Real-time data streaming and event-driven pipelines (Correct answer)
- Data visualization
Correct answer: Real-time data streaming and event-driven pipelines
Apache Kafka is a distributed event streaming platform used for high-throughput, real-time data pipelines and stream processing.
Question 67: What is OLAP in business intelligence?
- Object-Level Analytics Platform
- Online Analytical Processing for multidimensional analysis (Correct answer)
- Online Linear Analytics Processing
- Operational Lookup and Aggregation Protocol
Correct answer: Online Analytical Processing for multidimensional analysis
OLAP (Online Analytical Processing) enables fast, multidimensional analysis of large datasets by organizing data into cubes for slicing, dicing, and drilling.
Question 68: What is a dimension table in a data warehouse?
- A table storing descriptive attributes used to filter and group fact data (Correct answer)
- A lookup table for database indexes
- A table storing numerical measures like sales amount
- A table containing raw transactional data
Correct answer: A table storing descriptive attributes used to filter and group fact data
Dimension tables contain descriptive attributes (like customer name, product category, or date) used to provide context and filtering for measures in fact tables.
Question 69: Which SQL clause specifies the condition for filtering rows in a SELECT statement?
- FROM
- WHERE (Correct answer)
- SELECT
- GROUP BY
Correct answer: WHERE
The WHERE clause filters individual rows based on a specified condition before any grouping or aggregation occurs.
Question 70: Which statistical test compares means across more than two independent groups?
- Pearson correlation
- ANOVA (Correct answer)
- t-test
- Chi-square test
Correct answer: ANOVA
ANOVA (Analysis of Variance) tests whether the means of three or more independent groups are significantly different.
Question 71: What are the five dimensions commonly used to measure data quality?
- Availability, Reliability, Validity, Visibility, Volume
- Accuracy, Completeness, Consistency, Timeliness, Uniqueness (Correct answer)
- Format, Frequency, Freshness, Fidelity, Flexibility
- Speed, Scale, Security, Scope, Structure
Correct answer: Accuracy, Completeness, Consistency, Timeliness, Uniqueness
The five core data quality dimensions are Accuracy (correct values), Completeness (no missing data), Consistency (no contradictions), Timeliness (up to date), and Uniqueness (no duplicates).
Question 72: What is a streaming analytics pipeline?
- Real-time processing of continuous data streams as events arrive (Correct answer)
- A batch job that runs every hour
- A dashboard that refreshes daily
- A type of ETL process for historical data
Correct answer: Real-time processing of continuous data streams as events arrive
A streaming analytics pipeline processes data continuously as it arrives, enabling real-time insights rather than waiting for batch processing windows.
Question 73: What is the purpose of a validation set in machine learning?
- To augment the training data
- To train the model's parameters
- To provide a final unbiased evaluation of the model
- To tune hyperparameters and select the best model during development (Correct answer)
Correct answer: To tune hyperparameters and select the best model during development
A validation set is used during model development to tune hyperparameters and compare different models, separate from the test set used for final evaluation.
Question 74: What is the mean of the dataset: 4, 8, 6, 10, 12?
- 9
- 8 (Correct answer)
- 10
- 7
Correct answer: 8
The mean is calculated as (4+8+6+10+12)/5 = 40/5 = 8.
Question 75: What is a foreign key in a relational database?
- A key that stores NULL values
- An index on multiple columns
- A key used for encryption
- A column that references the primary key of another table (Correct answer)
Correct answer: A column that references the primary key of another table
A foreign key establishes a referential link between a column in one table and the primary key of another table.
Question 76: What is the Central Limit Theorem (CLT)?
- Larger datasets always have smaller standard deviations
- All populations follow a normal distribution
- The sampling distribution of the mean approaches normal as sample size increases (Correct answer)
- The mean always equals the median in a dataset
Correct answer: The sampling distribution of the mean approaches normal as sample size increases
The CLT states that the sampling distribution of the sample mean approximates a normal distribution as sample size increases, regardless of population distribution.
Question 77: Which SQL window function assigns a unique rank to each row within a partition?
- ROW_NUMBER() (Correct answer)
- PARTITION BY
- SUM() OVER
- LAG()
Correct answer: ROW_NUMBER()
ROW_NUMBER() assigns a unique sequential integer to each row within a partition, with no gaps or ties.
Question 78: Which Python library is most commonly used for statistical data visualization?
- NumPy
- Scikit-learn
- Seaborn (Correct answer)
- Pandas
Correct answer: Seaborn
Seaborn is built on Matplotlib and provides a high-level interface for drawing attractive and informative statistical graphics.
Question 79: What is the main purpose of a dashboard in data analytics?
- To store raw data efficiently
- To replace detailed analytical reports
- To provide a real-time overview of key metrics in one view (Correct answer)
- To automate data collection
Correct answer: To provide a real-time overview of key metrics in one view
A dashboard consolidates key performance indicators and metrics into a single, at-a-glance interface for monitoring business health.
Question 80: What does a PRIMARY KEY constraint enforce in a relational database?
- Default column values
- Referential integrity
- Unique, non-null values per row (Correct answer)
- Column data types
Correct answer: Unique, non-null values per row
A PRIMARY KEY ensures every row has a unique, non-null identifier in that column or set of columns.
Question 81: What is Snowflake's multi-cluster, shared data architecture?
- Separate storage and compute layers allowing multiple independent query engines (Correct answer)
- A single cluster serving all workloads
- Multiple databases sharing one compute cluster
- A peer-to-peer distributed computing model
Correct answer: Separate storage and compute layers allowing multiple independent query engines
Snowflake separates storage from compute, allowing multiple virtual warehouses to independently access the same data without contention.
Question 82: What is data lineage?
- The origin, movement, transformation, and destination of data throughout its lifecycle (Correct answer)
- A list of data owners and stewards
- The age of a dataset
- The organizational hierarchy of data systems
Correct answer: The origin, movement, transformation, and destination of data throughout its lifecycle
Data lineage tracks where data comes from, how it moves through systems, what transformations it undergoes, and where it ends up.
Question 83: What is a data lake?
- A relational database optimized for analytics
- A cloud-based ETL pipeline
- A centralized repository storing raw data in native format at any scale (Correct answer)
- A data warehouse with real-time capabilities
Correct answer: A centralized repository storing raw data in native format at any scale
A data lake stores structured, semi-structured, and unstructured data in its raw format, enabling flexible analysis without predefined schemas.
Question 84: Which chart type is best for showing the distribution of a single continuous variable?
- Line chart
- Bar chart
- Pie chart
- Histogram (Correct answer)
Correct answer: Histogram
A histogram displays the frequency distribution of continuous data by grouping values into bins.
Question 85: What is the purpose of data profiling?
- Analyzing datasets to understand structure, content, quality, and relationships (Correct answer)
- Building user profiles for personalization
- Profiling analyst behavior for security monitoring
- Creating performance profiles for databases
Correct answer: Analyzing datasets to understand structure, content, quality, and relationships
Data profiling examines dataset characteristics including structure, content, quality, completeness, and statistical properties to understand the data before use.
Question 86: What are the 3 Vs traditionally used to define Big Data?
- Variety, Veracity, Value
- Velocity, Volume, Variety (Correct answer)
- Value, Veracity, Velocity
- Volume, Visibility, Validity
Correct answer: Velocity, Volume, Variety
The original three defining characteristics of Big Data are Volume (size), Velocity (speed of generation), and Variety (different data types).
Question 87: What is Delta Lake?
- A real-time streaming platform
- A columnar file format for Hadoop
- A cloud storage service by AWS
- An open-source storage layer that brings ACID transactions to data lakes (Correct answer)
Correct answer: An open-source storage layer that brings ACID transactions to data lakes
Delta Lake is an open-source storage layer that adds ACID transaction support, schema enforcement, and time travel capabilities to data lakes.
Question 88: What is 'chart junk' in data visualization?
- Broken or corrupted chart files
- Charts with too many data points
- Incorrectly formatted axis labels
- Visual elements that add clutter without conveying information (Correct answer)
Correct answer: Visual elements that add clutter without conveying information
Chart junk refers to unnecessary visual elements like decorative patterns, 3D effects, and gridlines that clutter charts without adding meaning.
Question 89: What chart type is most appropriate for comparing parts of a whole across multiple categories?
- Stacked bar chart (Correct answer)
- Box plot
- Waterfall chart
- Scatter plot
Correct answer: Stacked bar chart
A stacked bar chart shows how each category's total is composed of subcategories, enabling part-to-whole comparisons.
Question 90: What type of correlation does a Pearson coefficient of -0.85 indicate?
- Strong negative correlation (Correct answer)
- No correlation
- Strong positive correlation
- Weak positive correlation
Correct answer: Strong negative correlation
A Pearson coefficient of -0.85 indicates a strong negative linear relationship where one variable increases as the other decreases.
Question 91: What is multicollinearity in regression analysis?
- Non-linear relationships between variables
- Unequal variance across residuals
- High correlation among independent predictor variables (Correct answer)
- Using multiple outcome variables
Correct answer: High correlation among independent predictor variables
Multicollinearity occurs when two or more independent variables in a regression model are highly correlated, making it difficult to isolate individual effects.
Question 92: What does 'data sovereignty' mean?
- The concept that data is subject to the laws of the country where it is stored or collected (Correct answer)
- The authority of the data governance team over data policies
- An individual's right to control their personal data
- A company's ownership of all its data assets
Correct answer: The concept that data is subject to the laws of the country where it is stored or collected
Data sovereignty means data is governed by the laws and regulations of the country in which it is physically located or collected.
CompTIA Data+ (DA0-001)
CompTIA Data+ validates skills in data mining, analysis, visualization, and governance for early-career data and business intelligence professionals. It covers the full data lifecycle from data concepts and mining through analysis, visualization, and governance quality controls.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds