Spark Core and RDDs Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Spark Core and RDDs flashcards as text
What does RDD stand for in Apache Spark?
Answer: Resilient Distributed Dataset
RDD stands for Resilient Distributed Dataset, which is the fundamental data structure in Apache Spark.
Which of the following is a transformation in Spark RDDs?
Answer: map()
map() is a transformation that returns a new RDD, while collect(), count(), and save() are actions.
What is the default number of partitions when creating an RDD from a collection using sc.parallelize()?
Answer: Depends on the SparkContext default parallelism
By default, sc.parallelize() uses the SparkContext's default parallelism, which is typically the number of cores.
Which RDD operation returns all elements of the RDD to the driver program?
Answer: collect()
collect() returns all elements of the RDD to the driver program as an array.
What does the flatMap() transformation do in Spark?
Answer: Maps each element to zero or more output elements and flattens the result
flatMap() applies a function to each element that returns an iterator and flattens all iterators into a single RDD.
Which storage level in Spark caches RDD data in memory as deserialized Java objects?
Answer: MEMORY_ONLY
MEMORY_ONLY stores RDD data in memory as deserialized Java objects, offering the fastest access but highest memory usage.