Data Ingestion Patterns Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Ingestion Patterns flashcards as text
An ingestion job must guarantee that a network retry does not double-insert rows. Which database technique most directly enables this?
Answer: Upsert/merge keyed on a unique business key
An upsert/merge on a unique key updates existing rows instead of inserting duplicates, making retries safe.
What is the purpose of a staging (landing) zone in an ingestion pipeline?
Answer: Hold raw ingested data before validation and transformation
A staging/landing zone temporarily holds raw data so it can be validated and transformed before loading downstream.
Which metric is most important for detecting that a real-time ingestion stream is falling behind?
Answer: Consumer lag
Consumer lag measures how far behind the consumer is from the latest offset, directly indicating a backlog.
To make a batch ingestion job safely re-runnable after a partial failure, the job should be:
Answer: Idempotent, often using partition overwrite for the target window
An idempotent job (e.g., overwriting the affected partition) can be safely re-run without producing duplicates.
When pulling from a rate-limited API, which strategy best handles HTTP 429 responses?
Answer: Exponential backoff with retry
Exponential backoff progressively increases wait time between retries, respecting rate limits and reducing pressure.
What is the main benefit of capturing data lineage during ingestion?
Answer: Tracing where each dataset came from for debugging and compliance
Data lineage records the origin and transformations of data, aiding debugging, auditing, and regulatory compliance.
In a fan-out ingestion pattern, a single source stream is:
Answer: Delivered to multiple independent consumers
Fan-out distributes one source stream to many independent consumers, each processing the data for its own purpose.