← All Data Engineering Flashcard Decks

Optimizing Query Performance Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Optimizing Query Performance flashcards as text
  1. Data skew in a distributed join causes which symptom?

    Answer: A few tasks run far longer than the rest

    Skew concentrates rows on a few keys, overloading the tasks that process them.

  2. Which technique mitigates skew on a hot join key?

    Answer: Salting the key to spread rows across partitions

    Salting appends a random component to distribute a hot key over more partitions.

  3. A covering index improves a query because it:

    Answer: Contains all columns the query needs, avoiding table lookups

    When the index includes every referenced column, the engine answers from the index alone.

  4. SELECT * in an analytical query on columnar storage is discouraged mainly because it:

    Answer: Forces reading all columns, increasing I/O

    Selecting every column negates columnar's ability to skip unused columns.

  5. Materializing an expensive, frequently-run aggregation into a summary table is best described as:

    Answer: A materialized view / precomputation tradeoff

    Precomputing results trades storage and refresh cost for faster reads.

  6. Which condition typically prevents an index from being used on a column?

    Answer: Wrapping the column in a function in the WHERE clause

    Functions applied to a column make the predicate non-sargable, blocking index use.

  7. When two large tables are joined, which strategy avoids broadcasting and instead shuffles both sides by key?

    Answer: Sort-merge / shuffle hash join

    For two large inputs, both are repartitioned by the join key and merged or hashed.