โ† All Data Engineering Flashcard Decks

Data Ingestion Patterns Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Ingestion Patterns flashcards as text
  1. An ingestion job must guarantee that a network retry does not double-insert rows. Which database technique most directly enables this?

    Answer: Upsert/merge keyed on a unique business key

    An upsert/merge on a unique key updates existing rows instead of inserting duplicates, making retries safe.

  2. What is the purpose of a staging (landing) zone in an ingestion pipeline?

    Answer: Hold raw ingested data before validation and transformation

    A staging/landing zone temporarily holds raw data so it can be validated and transformed before loading downstream.

  3. Which metric is most important for detecting that a real-time ingestion stream is falling behind?

    Answer: Consumer lag

    Consumer lag measures how far behind the consumer is from the latest offset, directly indicating a backlog.

  4. To make a batch ingestion job safely re-runnable after a partial failure, the job should be:

    Answer: Idempotent, often using partition overwrite for the target window

    An idempotent job (e.g., overwriting the affected partition) can be safely re-run without producing duplicates.

  5. When pulling from a rate-limited API, which strategy best handles HTTP 429 responses?

    Answer: Exponential backoff with retry

    Exponential backoff progressively increases wait time between retries, respecting rate limits and reducing pressure.

  6. What is the main benefit of capturing data lineage during ingestion?

    Answer: Tracing where each dataset came from for debugging and compliance

    Data lineage records the origin and transformations of data, aiding debugging, auditing, and regulatory compliance.

  7. In a fan-out ingestion pattern, a single source stream is:

    Answer: Delivered to multiple independent consumers

    Fan-out distributes one source stream to many independent consumers, each processing the data for its own purpose.