Introduction To Data Engineering Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Introduction To Data Engineering flashcards as text
What is the primary role of a data engineer in an organization?
Answer: Building and maintaining data pipelines and infrastructure
Data engineers build and maintain the pipelines and infrastructure that make data available for analysis.
Which process describes Extract, Transform, Load (ETL)?
Answer: Pulling data from sources, reshaping it, then storing it in a target system
ETL extracts data from sources, transforms it into a usable shape, and loads it into a destination.
What distinguishes a data warehouse from a data lake?
Answer: A warehouse stores structured, schema-defined data while a lake stores raw data of any format
Warehouses hold structured, schema-on-write data, while lakes hold raw data in many formats with schema-on-read.
Which of the following is a common batch processing framework?
Answer: Apache Spark
Apache Spark is a widely used distributed engine for batch (and stream) data processing.
What does 'data normalization' aim to reduce in a relational database?
Answer: Data redundancy and update anomalies
Normalization organizes tables to minimize redundancy and avoid anomalies during updates.
Which term describes data processed continuously as it arrives?
Answer: Stream processing
Stream processing handles data record-by-record in near real time as it arrives.
What is the purpose of a schema in a database?
Answer: To define the structure, types, and relationships of the data
A schema defines how data is organized, including tables, columns, types, and relationships.