← All Apache Spark Flashcard Decks

Spark Architecture and Cluster Management Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Spark Architecture and Cluster Management flashcards as text
  1. What is the purpose of spark.default.parallelism?

    Answer: Sets the default number of RDD partitions for transformations that create new partitions

    spark.default.parallelism sets the default number of partitions for RDD operations like reduceByKey and join when no explicit partition count is given.

  2. What is the role of the DAGScheduler in Apache Spark?

    Answer: Converts the RDD DAG into stages and submits task sets to the TaskScheduler

    The DAGScheduler transforms the logical DAG of RDD operations into physical stages and submits sets of tasks to the TaskScheduler for execution.

  3. What is the purpose of the BlockManager in Apache Spark?

    Answer: Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk

    BlockManager is Spark's distributed storage system that manages how data (RDD partitions, shuffle blocks, broadcasts) is stored across memory and disk.

  4. What is a Spark application's relationship to Spark jobs?

    Answer: An application contains one or more jobs, each triggered by an action

    A Spark application can trigger multiple jobs; each action (collect, count, save) creates one job consisting of stages and tasks.

  5. In which deploy mode does the Spark driver run on the machine where spark-submit is executed?

    Answer: Client mode

    In client mode, the Driver runs on the machine that submitted the application, which is useful for interactive development but not recommended for production.

  6. What is an accumulator in Apache Spark?

    Answer: A write-only distributed counter or aggregator that executors can add to and the driver can read

    Accumulators are write-only from executor perspective — tasks can add to them, but only the driver can read the accumulated value.