โ† All Data Engineering Flashcard Decks

Orchestrating Data Workflows Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Orchestrating Data Workflows flashcards as text
  1. What is the main advantage of defining workflows as code (e.g., Python DAGs) over a GUI-only tool?

    Answer: Version control, code review, and reproducibility

    Workflows-as-code can be versioned, reviewed, and tested like any other software.

  2. In Airflow, what does the schedule_interval '@daily' do?

    Answer: Runs the DAG once at midnight each day

    '@daily' schedules the DAG to run once per day at midnight.

  3. Which retry strategy helps avoid overwhelming a downstream system after repeated failures?

    Answer: Exponential backoff

    Exponential backoff increases the delay between retries, reducing load on failing systems.

  4. What problem does a task dependency graph primarily prevent?

    Answer: Running tasks before their inputs are ready

    Dependency graphs ensure a task only runs after its upstream prerequisites complete.

  5. In Prefect or Dagster, what is a key advantage over a pure cron schedule?

    Answer: Built-in retries, observability, and dependency management

    Modern orchestrators add observability, retries, and dependency handling that cron lacks.

  6. What does 'idempotency key' commonly guard against in pipeline writes?

    Answer: Duplicate records from re-processing the same batch

    An idempotency key lets the system detect and skip duplicate writes during retries.

  7. Which scenario is best handled by event-driven orchestration rather than time-based scheduling?

    Answer: Triggering a pipeline the moment a new file lands in cloud storage

    Event-driven orchestration reacts to events like file arrival rather than waiting for a clock.