โ† All Data Engineering Flashcard Decks

Data Quality and Testing Flashcards

6 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Quality and Testing flashcards as text
  1. Which metric measures the percentage of required fields that are populated with non-null values?

    Answer: Completeness rate

    Completeness rate measures what percentage of expected data fields actually contain values, identifying missing data problems.

  2. What is 'schema evolution' and why does it pose a challenge for data pipelines?

    Answer: Changes to the structure of data (added/removed/renamed columns) that can break downstream consumers

    Schema evolution occurs when source data structures change, potentially breaking downstream pipelines that expect a fixed schema.

  3. In the context of data testing, what is a 'unit test' for a dbt model?

    Answer: A test that validates a transformation's SQL logic using mocked input data

    A dbt unit test (introduced in dbt v1.8) runs the model's SQL against a small set of mocked input data to verify transformation logic in isolation.

  4. What is 'idempotency' in the context of data pipeline testing?

    Answer: The property that running a pipeline multiple times produces the same result as running it once

    An idempotent pipeline can be safely re-run multiple times without producing duplicate data or incorrect results.

  5. Which data observability tool monitors data pipelines by tracking metadata like row counts, null rates, and schema changes without accessing raw data?

    Answer: Monte Carlo

    Monte Carlo is a data observability platform that detects data quality incidents by monitoring pipeline metadata and statistical properties.

  6. What is the purpose of a 'data quality scorecard' in enterprise data engineering?

    Answer: To provide a consolidated view of data quality metrics across datasets, pipelines, and business domains

    A data quality scorecard aggregates quality metrics (completeness, accuracy, freshness) across datasets to give stakeholders a unified quality view.