← All Apache Spark Flashcard Decks

Spark Performance Tuning Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Spark Performance Tuning flashcards as text
  1. What is the purpose of the Spark Web UI?

    Answer: Monitoring running and completed jobs, stages, tasks, and executor metrics

    The Spark Web UI (default port 4040) provides detailed metrics on jobs, stages, tasks, storage, and executors for debugging and optimization.

  2. What happens when Spark runs out of memory for a shuffle and cannot fit data in memory?

    Answer: Spark spills the excess data to disk

    When execution memory is exhausted during a shuffle, Spark spills data to disk, which is slower but allows the job to continue.

  3. What is dynamic partition pruning in Spark 3.x?

    Answer: Filtering partitions of a fact table at runtime using values from a dimension table join

    Dynamic partition pruning pushes the result of a dimension table filter into the fact table scan, skipping irrelevant partitions at runtime.

  4. What is the difference between cache() and persist() in Spark?

    Answer: cache() is a shorthand for persist(MEMORY_ONLY); persist() allows specifying the storage level

    cache() is equivalent to persist(StorageLevel.MEMORY_ONLY), while persist() allows you to choose a different storage level.

  5. Which method forces Spark to recompute an RDD and remove its cached data?

    Answer: rdd.unpersist()

    unpersist() removes the RDD from the cache, freeing the memory for other computations.

  6. What is the recommended way to handle small file problems in Spark?

    Answer: Use coalesce() or repartition() to merge small partitions before writing

    Using coalesce() to reduce partitions before writing merges small files into larger ones, improving subsequent read performance.