Spark Core and RDDs Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Spark Core and RDDs flashcards as text
What is the purpose of the reduceByKey() transformation in Spark?
Answer: Aggregates values for each key using a given function
reduceByKey() merges values for each key using an associative and commutative function, performing a local combine before shuffling.
Which of these creates an RDD from an existing Scala/Python collection in Spark?
Answer: sc.parallelize()
sc.parallelize() creates an RDD by distributing an existing local collection across the cluster.
What is the role of the SparkContext in a Spark application?
Answer: Serves as the entry point for Spark functionality and represents the connection to a Spark cluster
SparkContext is the main entry point for Spark Core API, representing the connection to a Spark cluster and used to create RDDs.
What does the coalesce() transformation do in Spark?
Answer: Reduces the number of partitions with minimal data shuffling
coalesce() reduces the number of partitions by merging them, avoiding a full shuffle when decreasing partition count.
Which action computes the sum of all elements in a numeric RDD?
Answer: reduce(lambda a,b: a+b)
reduce() with a lambda that adds two values computes the sum across all elements in the RDD.
What does Spark's union() transformation return?
Answer: An RDD containing all elements from both RDDs, including duplicates
union() returns a new RDD containing all elements from both input RDDs, preserving duplicates.