Spark Performance Tuning Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Spark Performance Tuning flashcards as text
What is the benefit of using Parquet over CSV in Spark?
Answer: Parquet is columnar, enabling predicate pushdown and efficient column pruning
Parquet's columnar storage enables Spark to read only needed columns and push filters to the storage layer, drastically reducing I/O.
What is the Tungsten execution engine in Spark?
Answer: A low-level execution engine that uses off-heap memory and code generation for performance
Tungsten is Spark's physical execution engine that uses off-heap memory management and whole-stage code generation to maximize performance.
Which Spark configuration sets the amount of memory allocated per executor?
Answer: spark.executor.memory
spark.executor.memory sets the amount of memory (e.g., '4g') allocated to each Spark executor JVM process.
What is predicate pushdown in Spark SQL?
Answer: Moving filter conditions to earlier stages of the query plan to reduce data processed
Predicate pushdown moves filter conditions as close to the data source as possible, reducing the amount of data read from disk.
What does the spark.memory.fraction configuration control?
Answer: The fraction of JVM heap used for Spark's execution and storage memory combined
spark.memory.fraction (default 0.6) defines the fraction of JVM heap available for Spark's unified execution and storage memory pool.
What is whole-stage code generation in Spark?
Answer: Compiling an entire pipeline of operators into a single Java function to reduce interpretation overhead
Whole-stage code generation compiles multiple query operators into a single optimized function, eliminating virtual function calls and improving CPU efficiency.