โ† All Data Engineering Flashcard Decks

Fundamentals Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Fundamentals flashcards as text
  1. What is the primary difference between ETL and ELT?

    Answer: ETL transforms data before loading into the target; ELT loads raw data first then transforms inside the target

    In ETL transformation happens before loading, while in ELT raw data is loaded first and transformed within the destination system.

  2. Which file format is columnar and optimized for analytical query performance?

    Answer: Parquet

    Parquet stores data by column, enabling efficient compression and fast analytical reads.

  3. What does idempotency mean for a data pipeline task?

    Answer: Running the task multiple times produces the same result as running it once

    An idempotent task yields the same outcome no matter how many times it is executed, preventing duplicate side effects.

  4. In a star schema, the central table that holds measurable business events is called the:

    Answer: Fact table

    The fact table sits at the center of a star schema and stores quantitative metrics referencing surrounding dimension tables.

  5. Which tool is most commonly used to orchestrate and schedule batch data workflows as DAGs?

    Answer: Apache Airflow

    Apache Airflow defines workflows as directed acyclic graphs (DAGs) and schedules their execution.

  6. What is the main purpose of data partitioning in a large table?

    Answer: To improve query performance and manageability by dividing data into segments

    Partitioning splits a large dataset into smaller segments so queries scan less data, improving performance.

  7. A schema-on-read approach is most characteristic of which storage system?

    Answer: A data lake

    Data lakes apply schema-on-read, storing raw data and interpreting structure only when the data is queried.