Spark Performance Tuning Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Spark Performance Tuning flashcards as text
What is data skew in Apache Spark and why is it a problem?
Answer: Uneven data distribution across partitions causing some tasks to take much longer than others
Data skew occurs when partitions have unequal sizes, making some tasks significantly slower and creating bottlenecks in the job.
What does the broadcast join hint do in Spark SQL?
Answer: Replicates a small DataFrame to all executors to avoid shuffling the large DataFrame
Broadcast join sends a copy of the smaller DataFrame to all executor nodes, eliminating the expensive shuffle for the larger DataFrame.
What is the purpose of the spark.sql.shuffle.partitions configuration?
Answer: Controls the number of partitions used for shuffle operations like joins and aggregations
spark.sql.shuffle.partitions controls the number of partitions created after a shuffle (default 200), affecting performance of joins and aggregations.
What is speculative execution in Apache Spark?
Answer: Running duplicate copies of slow tasks on other nodes to handle stragglers
Speculative execution launches duplicate copies of straggler tasks on other nodes; whichever finishes first provides the result.
Which of the following strategies helps avoid data skew in a join operation?
Answer: Salting the join key with a random prefix
Salting adds a random prefix to skewed keys, distributing hot keys across multiple partitions to avoid overloading a single task.
What does Adaptive Query Execution (AQE) do in Spark 3.x?
Answer: Re-optimizes query plans at runtime based on actual statistics collected during execution
AQE re-optimizes the query plan at runtime using actual partition statistics, enabling better join strategy selection and skew handling.