← All Apache Spark Flashcard Decks

Spark Streaming Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Spark Streaming flashcards as text
  1. What is checkpointing used for in Spark Streaming?

    Answer: To save the state and metadata of a streaming application for fault recovery

    Checkpointing saves the streaming application's state (offsets, aggregation state) to fault-tolerant storage for recovery after failures.

  2. Which join type is supported in Spark Structured Streaming for stream-to-static dataset joins?

    Answer: Inner, left outer, right outer, and full outer joins

    Stream-static joins in Structured Streaming support inner, left outer, right outer, and full outer joins.

  3. What is the Complete output mode in Spark Structured Streaming?

    Answer: Writes the entire updated result table to the sink after every trigger

    Complete mode writes the entire result table to the sink after every trigger, suitable for aggregations that update existing results.

  4. Which source reads data from Apache Kafka in Spark Structured Streaming?

    Answer: spark.readStream.format('kafka')

    spark.readStream.format('kafka') reads streaming data from Kafka topics using the Kafka connector for Structured Streaming.

  5. What does the foreachBatch() sink do in Structured Streaming?

    Answer: Processes each micro-batch as a static DataFrame with a user-defined function

    foreachBatch() allows you to apply arbitrary operations on each micro-batch DataFrame, useful for writing to non-native sinks.

  6. In Spark Structured Streaming, what does event time refer to?

    Answer: The time when the record was generated at the source

    Event time is the time embedded in the data itself, representing when the event actually occurred at the source.