โ† All Data Engineering Flashcard Decks

Distributed Data Processing Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Distributed Data Processing flashcards as text
  1. In a MapReduce job, what is the primary purpose of the shuffle phase?

    Answer: To group and transfer mapper output to reducers by key

    The shuffle phase sorts mapper output and routes records with the same key to the same reducer.

  2. What problem does data skew most directly cause in a distributed join?

    Answer: A few overloaded tasks slow the whole stage

    Skew concentrates many records on one key, making a few tasks far slower than the rest.

  3. Which Spark operation triggers actual computation rather than just building the DAG?

    Answer: count()

    count() is an action, while filter, map, and select are lazy transformations.

  4. A broadcast join is most appropriate when:

    Answer: One table is small enough to fit in each executor's memory

    Broadcasting a small table to every node avoids shuffling the large table.

  5. What does HDFS block replication primarily provide?

    Answer: Fault tolerance against node failure

    Multiple replicas let the cluster recover data if a DataNode fails.

  6. In Spark, what is a 'narrow' transformation?

    Answer: Each output partition depends on one input partition

    Narrow transformations like map require no data movement across partitions.

  7. Why is checkpointing useful in long-running streaming jobs?

    Answer: It persists state so the job can recover after failure

    Checkpoints save progress and state, enabling recovery without reprocessing everything.