Spark Architecture and Cluster Management Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Spark Architecture and Cluster Management flashcards as text
What is the purpose of spark.default.parallelism?
Answer: Sets the default number of RDD partitions for transformations that create new partitions
spark.default.parallelism sets the default number of partitions for RDD operations like reduceByKey and join when no explicit partition count is given.
What is the role of the DAGScheduler in Apache Spark?
Answer: Converts the RDD DAG into stages and submits task sets to the TaskScheduler
The DAGScheduler transforms the logical DAG of RDD operations into physical stages and submits sets of tasks to the TaskScheduler for execution.
What is the purpose of the BlockManager in Apache Spark?
Answer: Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk
BlockManager is Spark's distributed storage system that manages how data (RDD partitions, shuffle blocks, broadcasts) is stored across memory and disk.
What is a Spark application's relationship to Spark jobs?
Answer: An application contains one or more jobs, each triggered by an action
A Spark application can trigger multiple jobs; each action (collect, count, save) creates one job consisting of stages and tasks.
In which deploy mode does the Spark driver run on the machine where spark-submit is executed?
Answer: Client mode
In client mode, the Driver runs on the machine that submitted the application, which is useful for interactive development but not recommended for production.
What is an accumulator in Apache Spark?
Answer: A write-only distributed counter or aggregator that executors can add to and the driver can read
Accumulators are write-only from executor perspective — tasks can add to them, but only the driver can read the accumulated value.