Spark Core and RDDs Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Spark Core and RDDs flashcards as text
Which method is used to persist an RDD in Spark?
Answer: cache() or persist()
cache() and persist() are used to persist an RDD; cache() is a shorthand for persist(StorageLevel.MEMORY_ONLY).
What is a narrow transformation in Spark?
Answer: A transformation where each input partition contributes to only one output partition
Narrow transformations are those where each input partition maps to only one output partition, requiring no data shuffle.
Which of the following is a wide transformation (shuffle) in Spark?
Answer: groupByKey()
groupByKey() is a wide transformation because it requires shuffling data across partitions to group all values for each key.
What is lineage in Apache Spark?
Answer: The record of transformations used to build an RDD from base data
Lineage is the sequence of transformations that were applied to create an RDD, allowing Spark to recompute lost partitions.
Which Spark RDD action returns the first n elements of the RDD?
Answer: take(n)
take(n) returns the first n elements of the RDD to the driver program.
How does Spark achieve fault tolerance with RDDs?
Answer: By recomputing lost partitions using lineage information
Spark achieves fault tolerance by recomputing lost RDD partitions using the recorded lineage of transformations.