← All Databricks Certified Data Engineer Associate Flashcard Decks

Auto Loader and Data Ingestion Flashcards

7 cards from real Databricks Certified Data Engineer Associate practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Auto Loader and Data Ingestion flashcards as text
  1. Which statement BEST describes the COPY INTO SQL command in Databricks?

    Answer: It is an idempotent SQL command that loads files from cloud storage into a Delta table, automatically skipping already-loaded files

    COPY INTO is a Databricks SQL command that idempotently loads data from cloud object storage into a Delta table, tracking which files have already been loaded to avoid duplicates.

  2. When an Auto Loader stream is started with `trigger(availableNow=True)`, what behavior results?

    Answer: The stream processes all data available at startup across one or more micro-batches, then stops automatically

    `availableNow=True` causes the stream to process all currently available data in as many micro-batches as needed, then terminate — combining streaming efficiency with batch semantics.

  3. What does setting `cloudFiles.useNotifications` to true accomplish in Auto Loader?

    Answer: Switches Auto Loader from directory listing mode to cloud file notification mode using services like AWS SNS/SQS

    Setting `cloudFiles.useNotifications=true` switches Auto Loader to file notification mode, where it provisions and listens to cloud-native event services (SNS/SQS on AWS) instead of polling the directory.

  4. In a production Auto Loader deployment, what information is stored at the `checkpointLocation`?

    Answer: Stream metadata, progress offsets, and the complete record of processed source files

    The checkpoint directory stores Structured Streaming metadata including offsets, committed file lists, and other state needed to resume processing exactly where it left off after a failure.

  5. Which of the following is the correct syntax to create an Auto Loader stream reading JSON files from a cloud path?

    Answer: spark.readStream.format('cloudFiles').option('cloudFiles.format', 'json').load(path)

    Auto Loader uses `format('cloudFiles')` as the source format and `cloudFiles.format` to specify the underlying file format — in this case, `json`.

  6. A data engineer needs to ingest millions of small JSON files uploaded daily to S3. Which approach is MOST appropriate?

    Answer: Auto Loader with file notification mode and cloudFiles.maxFilesPerTrigger to control batch size

    Auto Loader with file notification mode is optimal for high-volume, continuously arriving small files because it avoids full directory scans and processes only new files efficiently.

  7. When Auto Loader's `schemaEvolutionMode` is set to `'addNewColumns'` and a new column appears in incoming data, what occurs?

    Answer: The stream stops, updates the persisted schema at schemaLocation to include the new column, then restarts automatically

    With `addNewColumns` mode, Auto Loader detects the new column, updates the schema stored at `schemaLocation`, and automatically restarts the stream with the evolved schema to include the new column.