← All Apache Spark Flashcard Decks

Spark Core and RDDs Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Spark Core and RDDs flashcards as text
  1. What is the purpose of the reduceByKey() transformation in Spark?

    Answer: Aggregates values for each key using a given function

    reduceByKey() merges values for each key using an associative and commutative function, performing a local combine before shuffling.

  2. Which of these creates an RDD from an existing Scala/Python collection in Spark?

    Answer: sc.parallelize()

    sc.parallelize() creates an RDD by distributing an existing local collection across the cluster.

  3. What is the role of the SparkContext in a Spark application?

    Answer: Serves as the entry point for Spark functionality and represents the connection to a Spark cluster

    SparkContext is the main entry point for Spark Core API, representing the connection to a Spark cluster and used to create RDDs.

  4. What does the coalesce() transformation do in Spark?

    Answer: Reduces the number of partitions with minimal data shuffling

    coalesce() reduces the number of partitions by merging them, avoiding a full shuffle when decreasing partition count.

  5. Which action computes the sum of all elements in a numeric RDD?

    Answer: reduce(lambda a,b: a+b)

    reduce() with a lambda that adds two values computes the sum across all elements in the RDD.

  6. What does Spark's union() transformation return?

    Answer: An RDD containing all elements from both RDDs, including duplicates

    union() returns a new RDD containing all elements from both input RDDs, preserving duplicates.