AWS Certified Data Engineer – Associate (DEA-C01) — Questions and Answers
Question 1: In a workflow scheduler, what is 'backfilling'?
- Compressing intermediate data
- Running a DAG for past dates that were missed or newly added (Correct answer)
- Increasing worker memory
- Deleting old task logs
Correct answer: Running a DAG for past dates that were missed or newly added
Backfilling executes a pipeline over historical intervals to populate missing data.
Question 2: What is a 'task instance' in Airflow?
- A specific run of a task for a particular execution date (Correct answer)
- A database connection
- A worker node
- A reusable operator template
Correct answer: A specific run of a task for a particular execution date
A task instance is one execution of a task tied to a specific DAG run and date.
Question 3: Which format provides the best compression ratios for analytical data due to columnar storage and value encoding techniques?
- XML
- Parquet (Correct answer)
- JSON
- CSV
Correct answer: Parquet
Parquet achieves excellent compression by storing similar data types together in columns, enabling encodings like dictionary encoding and run-length encoding that exploit data locality.
Question 4: What does 'idempotency key' commonly guard against in pipeline writes?
- Slow queries
- Unencrypted data
- Missing schemas
- Duplicate records from re-processing the same batch (Correct answer)
Correct answer: Duplicate records from re-processing the same batch
An idempotency key lets the system detect and skip duplicate writes during retries.
Question 5: Which process describes Extract, Transform, Load (ETL)?
- Encrypting data before transmission
- Loading raw data first, then querying it directly
- Compressing files for backup
- Pulling data from sources, reshaping it, then storing it in a target system (Correct answer)
Correct answer: Pulling data from sources, reshaping it, then storing it in a target system
ETL extracts data from sources, transforms it into a usable shape, and loads it into a destination.
Question 6: What is the primary difference between ETL and ELT?
- ETL is only for streaming data while ELT is only for batch
- ETL requires a data lake while ELT requires a relational database
- ETL transforms data before loading into the target; ELT loads raw data first then transforms inside the target (Correct answer)
- There is no difference; the terms are interchangeable
Correct answer: ETL transforms data before loading into the target; ELT loads raw data first then transforms inside the target
In ETL transformation happens before loading, while in ELT raw data is loaded first and transformed within the destination system.
Question 7: What does a privacy-by-design approach emphasize?
- Building privacy protections into systems from the start (Correct answer)
- Disabling all logging
- Maximizing data collection
- Adding privacy only after a breach
Correct answer: Building privacy protections into systems from the start
Privacy by design embeds protections into system design upfront.
Question 8: What is a common reason to use a sensor in an orchestration tool?
- To compress output files
- To encrypt connection strings
- To wait for an external condition like a file or partition to appear (Correct answer)
- To allocate GPU resources
Correct answer: To wait for an external condition like a file or partition to appear
Sensors pause a workflow until a specified external event or condition is met.
Question 9: Which factor most directly limits the scalability of a global ORDER BY in distributed SQL?
- All data must funnel through a single sort/merge step (Correct answer)
- Lack of partitions
- Too many executors
- Columnar storage
Correct answer: All data must funnel through a single sort/merge step
Total ordering forces a global merge that becomes a bottleneck at scale.
Question 10: When ingesting from a REST API that paginates results, what mechanism prevents missing or duplicating records across pages?
- A stable cursor or continuation token (Correct answer)
- Random page ordering
- Increasing the request timeout
- Disabling retries
Correct answer: A stable cursor or continuation token
A stable cursor/continuation token gives a consistent position so pages neither overlap nor skip records.
Question 11: Many modern cloud data warehouses utilize a columnar storage format. What is the primary advantage of this format for analytical workloads?
- It improves the speed of transactional writes and updates on individual rows.
- It simplifies data ingestion from row-based sources like relational databases.
- It guarantees ACID compliance for all operations within the data warehouse.
- It allows for highly efficient data compression and reduces I/O by only reading the columns required for a query. (Correct answer)
Correct answer: It allows for highly efficient data compression and reduces I/O by only reading the columns required for a query.
Columnar storage organizes data by column rather than by row. This is highly advantageous for analytical queries, which typically only access a subset of columns. The system can read only the data from the required columns, significantly reducing I/O. Furthermore, since data within a column is of the same type, it can be compressed much more effectively than row-based data.
Question 12: Which approach minimizes small-file problems in a cloud data lake?
- Compacting many small files into larger files (Correct answer)
- Disabling partitioning entirely
- Writing one file per record
- Storing everything as JSON
Correct answer: Compacting many small files into larger files
Compaction merges many small files into fewer large files, improving read throughput and reducing overhead.
Question 13: A financial services company's compliance department is auditing a critical regulatory report and discovers a discrepancy in a key metric. To investigate, they need to trace the metric back through all the transformations and data pipelines to its original source systems. What data governance capability is essential for this investigation?
- Master Data Management (MDM)
- Data Lineage (Correct answer)
- Data Cataloging
- Data Encryption
Correct answer: Data Lineage
Data lineage provides a complete audit trail of data's journey, showing its origin, every transformation it undergoes, and its final destination. [1, 4, 6] This visibility is crucial for root cause analysis of errors, impact analysis of changes, and meeting regulatory compliance requirements by proving the provenance of data in reports. [5, 15]
Question 14: An analytics team frequently runs queries that aggregate metrics across a few specific columns (e.g., `SUM(revenue)`, `AVG(quantity)`) from a wide table with over 200 columns. The current storage format is row-based (like Avro). Queries are slow because they read a lot of unnecessary data. To improve performance for these specific analytical queries, which change would be most impactful?
- Denormalizing the data into an even wider table.
- Creating a B-tree index on every one of the 200 columns.
- Switching the storage format to a columnar format like Parquet or ORC. (Correct answer)
- Increasing the number of nodes in the processing cluster.
Correct answer: Switching the storage format to a columnar format like Parquet or ORC.
Columnar formats like Parquet or ORC store data by column instead of by row. When a query only needs to access a few columns, the query engine can read just the data for those specific columns, dramatically reducing I/O and improving performance for analytical workloads.
Question 15: Inmon's approach to data warehousing favors:
- Dimensional marts first, then a warehouse
- A normalized enterprise data warehouse feeding dependent marts (Correct answer)
- No central warehouse at all
- Storing only flat files
Correct answer: A normalized enterprise data warehouse feeding dependent marts
Inmon advocates a top-down, normalized (3NF) enterprise warehouse from which dependent data marts are derived.
Question 16: What is the purpose of a staging area in a data warehouse load process?
- To archive deleted records permanently
- To store user credentials
- To serve final reports to users
- To temporarily hold raw extracted data before transformation and loading (Correct answer)
Correct answer: To temporarily hold raw extracted data before transformation and loading
A staging area is intermediate storage where raw data lands before being cleaned, transformed, and loaded.
Question 17: In Apache Iceberg table format, what enables time-travel queries on large datasets?
- Bloom filter indexes
- Z-ordering of data files
- Immutable snapshot-based metadata and manifest files (Correct answer)
- Partition pruning metadata
Correct answer: Immutable snapshot-based metadata and manifest files
Iceberg uses immutable snapshots where each write creates a new snapshot, allowing queries to reference past snapshots for time-travel without data copies.
Question 18: Which column would you index to speed up 'current record' lookups in a Type 2 dimension?
- is_current flag (or end_date) (Correct answer)
- surrogate key alone
- explanation text column
- load_id
Correct answer: is_current flag (or end_date)
Filtering on is_current or the sentinel end_date is the common current-state access pattern.
Question 19: In Spark, persisting an RDD/DataFrame with caching mainly helps when:
- It is used only once
- It has no transformations
- It is written immediately to disk
- It is reused across multiple actions (Correct answer)
Correct answer: It is reused across multiple actions
Caching avoids recomputation when the same dataset feeds several actions.
Question 20: Why are surrogate keys preferred over natural keys when joining facts to a Type 2 dimension?
- They avoid the need for indexes
- They point a fact to the correct dimension version active at event time (Correct answer)
- They are always smaller integers
- They prevent NULL values in facts
Correct answer: They point a fact to the correct dimension version active at event time
The surrogate key resolves to the specific historical row valid when the fact occurred.
Question 21: In a star schema, the central table that holds measurable business events is called the:
- Lookup table
- Dimension table
- Fact table (Correct answer)
- Bridge table
Correct answer: Fact table
The fact table sits at the center of a star schema and stores quantitative metrics referencing surrounding dimension tables.
Question 22: In a galaxy (fact constellation) schema, what is shared across multiple fact tables?
- Conformed dimension tables (Correct answer)
- Surrogate keys only
- A single ETL job
- The same grain
Correct answer: Conformed dimension tables
A galaxy schema links multiple fact tables through shared conformed dimensions.
Question 23: Which technique helps prevent unauthorized data exfiltration from an organization?
- Data partitioning
- Query caching
- Data Loss Prevention (DLP) (Correct answer)
- Indexing
Correct answer: Data Loss Prevention (DLP)
DLP tools detect and block unauthorized movement of sensitive data.
Question 24: A large e-commerce platform runs its application on hundreds of virtual machines. The operations team needs to centralize all application logs in real-time for monitoring and security analysis. Which of the following represents the most common and effective architectural pattern for this task?
- Installing a lightweight agent (e.g., Fluentd, Filebeat) on each server to tail log files and stream events to a central log aggregator. (Correct answer)
- Writing a custom script on each server to periodically use SSH/SCP to copy log files to a central server.
- Configuring a cron job on each server to batch-compress and upload log files to cloud storage every hour.
- Directly writing logs from the application on each server to a central relational database.
Correct answer: Installing a lightweight agent (e.g., Fluentd, Filebeat) on each server to tail log files and stream events to a central log aggregator.
The standard and most robust pattern for centralized logging is to use a dedicated log shipping agent. Lightweight agents like Fluentd, Filebeat, or the OpenTelemetry Collector are designed specifically to run on each source machine, tail log files efficiently, and forward the log events in a streaming fashion to a central system (like Elasticsearch, OpenSearch, or a message queue). This approach is scalable, resilient to network issues, and provides real-time data. Hourly batch uploads are not real-time, direct database writes would cause a performance bottleneck, and custom SSH scripts are brittle and hard to manage at scale.
Question 25: Which approach best handles late-arriving data in a streaming pipeline?
- Switching to batch only
- Restarting the pipeline daily
- Ignoring all delayed records
- Using watermarks and windowing to account for event-time delays (Correct answer)
Correct answer: Using watermarks and windowing to account for event-time delays
Watermarks combined with event-time windows allow streaming systems to correctly handle records that arrive late.
Question 26: What does RBAC stand for in access control?
- Role-Based Access Control (Correct answer)
- Rule-Based Audit Compliance
- Realtime Batch Access Coordination
- Resource Backup And Caching
Correct answer: Role-Based Access Control
RBAC assigns permissions based on user roles.
Question 27: What is the purpose of an idempotent task in a data pipeline?
- It never needs to be scheduled
- It runs faster on each retry
- Re-running it produces the same result without side effects (Correct answer)
- It bypasses dependency checks
Correct answer: Re-running it produces the same result without side effects
Idempotent tasks can be safely retried because repeated execution yields the same end state.
Question 28: What is the main risk of trigger-based CDC compared to log-based CDC?
- Inability to detect deletes
- Added write overhead and latency on the source database (Correct answer)
- Producing only batch output
- Requiring a data warehouse
Correct answer: Added write overhead and latency on the source database
Trigger-based CDC fires triggers on every change, adding write overhead and latency to source transactions.
Question 29: A data engineering team is designing a data warehouse for a large financial institution. They need to prioritize storage efficiency and data integrity due to complex, multi-level hierarchies in their customer and account dimensions. Query performance is a secondary concern. Which schema design would be most appropriate for this scenario?
- Snowflake Schema (Correct answer)
- Star Schema
- Data Vault
- Flat Denormalized Model
Correct answer: Snowflake Schema
A Snowflake Schema is the most appropriate choice because it normalizes dimension tables into multiple related tables. This reduces data redundancy and improves data integrity, which is crucial for complex hierarchies. [1, 5] While this leads to more complex queries with more joins, it aligns with the stated priorities of storage efficiency and data integrity over query speed. [3, 6]
Question 30: Which process tracks the origin and transformations of data through a pipeline?
- Data deduplication
- Data sharding
- Data lineage (Correct answer)
- Data masking
Correct answer: Data lineage
Data lineage records data's source and how it changes over time.
Question 31: What does 'idempotency' mean for a data pipeline operation?
- It always doubles the data
- It can only run on weekends
- It deletes the source after running
- Running it multiple times produces the same result as running it once (Correct answer)
Correct answer: Running it multiple times produces the same result as running it once
An idempotent operation yields the same outcome no matter how many times it is executed.
Question 32: Which of the following features is a primary advantage of using a dedicated workflow orchestrator over a series of cron jobs for managing data pipelines?
- Native support for complex dependency management and failure handling with retries (Correct answer)
- Ability to execute scripts on a time-based schedule
- Simplified setup for running a single, isolated script
- Lower resource consumption for simple, single-task workflows
Correct answer: Native support for complex dependency management and failure handling with retries
While cron is excellent for simple time-based scheduling, it lacks built-in mechanisms for managing dependencies (e.g., 'run task B only if task A succeeds'), handling retries, backfilling, and providing centralized logging and monitoring, which are core features of orchestrators.
Question 33: Which of the following statements most accurately describes the primary difference between a data lake and a data warehouse in terms of data structure?
- A data lake stores raw data in its native format, applying structure during analysis ("schema-on-read"). (Correct answer)
- A data lake uses a "schema-on-write" approach, while a data warehouse uses "schema-on-read".
- A data warehouse stores data in a normalized form (3NF), while a data lake uses a denormalized star schema.
- A data lake stores only unstructured data, while a data warehouse stores only structured data.
Correct answer: A data lake stores raw data in its native format, applying structure during analysis ("schema-on-read").
The fundamental difference lies in when the schema is applied. A data warehouse requires a predefined schema before data is loaded (schema-on-write). In contrast, a data lake stores data in its raw, native format and the schema is applied when the data is read or queried for a specific analysis (schema-on-read), providing greater flexibility.
Question 34: A financial services company wants to monitor credit card transactions for fraudulent activity. The system must group all transactions by a specific user that occur closely together, but the time between these bursts of activity is unpredictable. A new group should start only after a significant period of user inactivity. Which windowing strategy is most suitable for this scenario?
- Global Windows
- Sliding Windows
- Tumbling Windows
- Session Windows (Correct answer)
Correct answer: Session Windows
Session windows are designed specifically for this use case. They group events based on periods of activity, which are terminated by a predefined gap of inactivity (a timeout). This allows the system to dynamically create windows for each user's transaction burst without having fixed start or end times.
Question 35: What does a data retention policy define?
- Which encryption algorithm to use
- How fast queries must run
- The number of replicas per table
- How long data is kept before deletion or archival (Correct answer)
Correct answer: How long data is kept before deletion or archival
Retention policies specify how long data is stored before disposal.
Question 36: Which storage option is best suited for a high-throughput streaming workload requiring low-latency appends?
- A managed message/log store like Kafka or Kinesis (Correct answer)
- Read replica of a relational DB
- Glacier archive
- Cold object storage
Correct answer: A managed message/log store like Kafka or Kinesis
Log-based streaming stores are designed for high-throughput, low-latency append operations.
Question 37: What is a common reason to use S3 Intelligent-Tiering?
- Automatic cost optimization for data with unpredictable access patterns (Correct answer)
- Guaranteed lowest latency
- Free data egress
- Built-in SQL querying
Correct answer: Automatic cost optimization for data with unpredictable access patterns
Intelligent-Tiering automatically moves objects between access tiers based on usage to optimize cost.
Question 38: What is the grain of a fact table?
- The total row count
- The number of dimensions attached
- The storage format used
- The level of detail represented by one row (Correct answer)
Correct answer: The level of detail represented by one row
Grain defines exactly what a single fact row represents, such as one line item per order.
Question 39: What is a data mart?
- A real-time streaming buffer
- A tool for writing SQL
- A backup of the entire data warehouse
- A subset of a data warehouse focused on a specific business area (Correct answer)
Correct answer: A subset of a data warehouse focused on a specific business area
A data mart is a focused, department-specific subset of a data warehouse serving a particular business function.
Question 40: An SCD Type 3 dimension tracks change by:
- Storing a limited 'previous value' column alongside the current value (Correct answer)
- Using a bridge table
- Adding a new row per change
- Overwriting silently
Correct answer: Storing a limited 'previous value' column alongside the current value
SCD Type 3 keeps a previous-value column, allowing comparison between current and one prior state.
Question 41: A data engineering team is building a data lake on a major cloud platform. The primary requirement is to store massive volumes of raw, semi-structured (JSON logs) and unstructured (images, videos) data in its native format. Which cloud storage solution is most appropriate for the foundational layer of this data lake?
- A managed relational database (e.g., Cloud SQL, RDS)
- An in-memory cache (e.g., Redis, Memcached)
- Block Storage (e.g., Amazon EBS, Google Persistent Disk)
- Object Storage (e.g., Amazon S3, Google Cloud Storage, Azure Blob Storage) (Correct answer)
Correct answer: Object Storage (e.g., Amazon S3, Google Cloud Storage, Azure Blob Storage)
Object storage is designed for storing vast amounts of unstructured and semi-structured data. It offers a flat namespace, high durability and availability, virtually limitless scalability, and a low cost per GB, making it the ideal foundation for a data lake where raw data is landed before processing.
Question 42: Why does idempotency matter in at-least-once processing?
- It speeds up the network
- It makes duplicate processing produce the same result (Correct answer)
- It guarantees ordering
- It removes partitions
Correct answer: It makes duplicate processing produce the same result
Idempotent operations let duplicate deliveries occur without corrupting results.
Question 43: What is the main benefit of data compression in storage systems?
- Increased password strength
- Better screen brightness
- Faster typing speed
- Reduced storage costs and faster I/O for large datasets (Correct answer)
Correct answer: Reduced storage costs and faster I/O for large datasets
Compression shrinks data size, lowering storage costs and often speeding up read/write operations.
Question 44: Which encryption option lets you manage your own keys for cloud-stored data while the provider performs encryption?
- Customer-managed keys via KMS (SSE-KMS) (Correct answer)
- No encryption
- Provider-managed keys (SSE-S3)
- Client-side hashing
Correct answer: Customer-managed keys via KMS (SSE-KMS)
Customer-managed keys through a key management service give control over key rotation and access policies.
Question 45: A bridge table is used to resolve which kind of relationship between facts and dimensions?
- Many-to-many (Correct answer)
- No relationship
- Self-referencing only
- One-to-one
Correct answer: Many-to-many
Bridge tables handle many-to-many relationships, such as a single bank account with multiple owners.
Question 46: Which US law protects the privacy of patient health information?
- CCPA
- GDPR
- GLBA
- HIPAA (Correct answer)
Correct answer: HIPAA
HIPAA governs protected health information in the US.
Question 47: What does horizontal scaling (scaling out) mean for a data system?
- Upgrading to faster disks only
- Reducing the dataset size
- Adding more machines or nodes to distribute the load (Correct answer)
- Adding more CPU and RAM to a single machine
Correct answer: Adding more machines or nodes to distribute the load
Horizontal scaling adds more nodes to share the workload, as opposed to upgrading a single machine (vertical scaling).
Question 48: Z-ordering / data clustering on commonly filtered columns helps by:
- Forcing a full scan
- Encrypting the data
- Removing partition keys
- Co-locating related values so more files can be skipped (Correct answer)
Correct answer: Co-locating related values so more files can be skipped
Clustering sorts related values together, improving data-skipping via file-level min/max stats.
Question 49: Which practice best supports recovering a pipeline from a mid-run failure?
- Deleting intermediate outputs immediately
- Checkpointing progress and making tasks resumable (Correct answer)
- Running everything in a single transaction with no logs
- Disabling retries
Correct answer: Checkpointing progress and making tasks resumable
Checkpointing and resumable tasks let a pipeline continue from where it failed rather than restarting fully.
Question 50: Which technique replaces sensitive values with non-sensitive surrogates that can be mapped back via a secure vault?
- Tokenization (Correct answer)
- Normalization
- Partitioning
- Caching
Correct answer: Tokenization
Tokenization substitutes data with tokens stored separately.
Question 51: What is the purpose of a Data Protection Officer (DPO) under GDPR?
- Oversee an organization's data protection strategy and compliance (Correct answer)
- Manage the database backup schedule
- Write ETL transformation code
- Optimize warehouse query performance
Correct answer: Oversee an organization's data protection strategy and compliance
A DPO ensures the organization complies with data protection laws.
Question 52: Why is a separate staging environment valuable when deploying pipeline changes?
- It lets you validate DAG changes before they affect production data (Correct answer)
- It runs faster than production
- It removes dependencies
- It eliminates the need for testing
Correct answer: It lets you validate DAG changes before they affect production data
Staging environments catch errors in workflow changes before they impact production.
Question 53: What is the primary advantage of using a wide table (denormalized) in analytical workloads?
- Fewer joins required at query time, improving performance (Correct answer)
- Better data integrity enforcement
- Easier schema evolution
- Reduced storage costs
Correct answer: Fewer joins required at query time, improving performance
Wide denormalized tables pre-join data, eliminating costly join operations at query time and improving analytical query performance.
Question 54: What is horizontal scaling (scaling out) in distributed data systems?
- Compressing the database
- Reducing the number of tables
- Upgrading a single machine's CPU and RAM
- Adding more machines to share the workload (Correct answer)
Correct answer: Adding more machines to share the workload
Horizontal scaling adds more nodes to distribute load, unlike vertical scaling which upgrades one machine.
Question 55: Which of the following best describes a data mart?
- A schema design that uses multiple fact tables sharing common dimensions.
- A subset of a data warehouse focused on a specific business line or department. (Correct answer)
- A large, centralized repository for all enterprise data from various sources.
- An un-modeled, schema-on-read repository for raw structured and unstructured data.
Correct answer: A subset of a data warehouse focused on a specific business line or department.
A data mart is a smaller, focused subset of a data warehouse that is designed for the specific needs of a particular department or business function, such as sales, finance, or marketing. [7, 9, 10] This allows for faster, more tailored access to relevant data for a specific group of users. [15]
Question 56: In the Kimball methodology, what is the 'Bus Matrix' used for?
- Defining the physical partitioning strategy
- Scheduling ETL job dependencies
- Documenting data lineage from source to target
- Mapping which business processes share which conformed dimensions (Correct answer)
Correct answer: Mapping which business processes share which conformed dimensions
The Enterprise Bus Matrix maps each business process (fact table) against the dimensions it shares, ensuring conformed dimensions across the data warehouse.
Question 57: Why is partitioning ingested data by event time often preferred over ingestion time for analytics?
- It keeps late-arriving events grouped with the period they actually belong to (Correct answer)
- It prevents all duplicates
- It guarantees faster writes
- It removes the need for a watermark
Correct answer: It keeps late-arriving events grouped with the period they actually belong to
Event-time partitioning places records in the time bucket they occurred, even if they arrive late, improving analytical correctness.
Question 58: Which encryption approach protects data while it is stored on disk in a data warehouse?
- Encryption at rest (Correct answer)
- Encryption in transit
- Token bucketing
- TLS handshake
Correct answer: Encryption at rest
Encryption at rest secures data persisted on storage media.
Question 59: What is the purpose of a manifest or metadata layer in formats like Delta Lake?
- To encrypt the data
- To replace the underlying object store
- To track which files belong to a table version for consistent reads (Correct answer)
- To compress raw files
Correct answer: To track which files belong to a table version for consistent reads
The metadata/transaction log records the set of files per snapshot, enabling ACID reads and time travel.
Question 60: What does 'run-length encoding' (RLE) do in the context of Parquet column compression?
- Stores the byte length of each value for fast random access
- Encrypts sensitive column values using a repeating key
- Replaces consecutive repeated values by storing the value and its repetition count (Correct answer)
- Encodes data by running multiple compression passes sequentially
Correct answer: Replaces consecutive repeated values by storing the value and its repetition count
Run-length encoding replaces consecutive repeated values with a single value and a count (e.g., five 'US' values become 'US×5'), which is very effective for sorted or low-cardinality columns.
Question 61: Which columns are most commonly added to support a Type 2 dimension?
- created_by and updated_by
- effective_date, end_date, and is_current flag (Correct answer)
- hash_key and load_id only
- primary_key and foreign_key
Correct answer: effective_date, end_date, and is_current flag
Validity dates plus a current-record indicator are the standard Type 2 tracking columns.
Question 62: What is the primary risk of a poorly designed long-running monolithic task in a DAG?
- It is too easy to test
- It uses no resources
- Failures require re-running everything with no granular recovery (Correct answer)
- It cannot be scheduled
Correct answer: Failures require re-running everything with no granular recovery
Monolithic tasks lack checkpoints, so any failure forces a full re-run.
Question 63: Which statement best describes the difference between scheduling and orchestration?
- Scheduling triggers jobs by time; orchestration manages dependencies and coordination (Correct answer)
- They are identical terms
- Orchestration only runs single tasks
- Scheduling handles error recovery exclusively
Correct answer: Scheduling triggers jobs by time; orchestration manages dependencies and coordination
Scheduling decides when jobs run, while orchestration coordinates dependencies, retries, and data flow across tasks.
Question 64: When implementing an SCD Type 2 dimension table for employees, a data engineer uses `effective_start_date` and `effective_end_date` columns to track the time period for which each record is valid. When an employee's department changes, which of the following actions must be performed?
- Delete the old record and insert a new record with the current date as the `effective_start_date`.
- Add a new column named `previous_department` and populate it with the old department name.
- Update the `department` field in the existing record and set the `effective_start_date` to the current date.
- Update the `effective_end_date` of the current record to the day before the change and insert a new record for the new department. (Correct answer)
Correct answer: Update the `effective_end_date` of the current record to the day before the change and insert a new record for the new department.
The standard procedure for managing SCD Type 2 with effective dates is to 'close' the currently active record by updating its `effective_end_date`. A new record is then inserted with the updated information, and its `effective_start_date` is set to the date the change became effective. This maintains a continuous and non-overlapping history.
Question 65: Which approach improves observability of a data pipeline?
- Removing all alerts
- Emitting metrics, structured logs, and lineage metadata (Correct answer)
- Disabling logs to save space
- Running everything as one task
Correct answer: Emitting metrics, structured logs, and lineage metadata
Metrics, structured logs, and lineage make pipeline behavior transparent and debuggable.
AWS Certified Data Engineer – Associate (DEA-C01)
The AWS Certified Data Engineer – Associate validates expertise in designing, building, and maintaining data pipelines and architectures on AWS. It covers data ingestion, transformation, storage management, operations, and security and governance.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds