Auto Loader and Data Ingestion Flashcards
7 cards from real Databricks Certified Data Engineer Associate practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Auto Loader and Data Ingestion flashcards as text
Which Auto Loader option limits processing to a maximum number of files per trigger interval?
Answer: cloudFiles.maxFilesPerTrigger
`cloudFiles.maxFilesPerTrigger` caps the number of files Auto Loader will process in a single micro-batch, useful for controlling throughput and cluster load.
Auto Loader's file notification mode on AWS relies on which combination of managed services?
Answer: S3 + SNS + SQS
On AWS, Auto Loader's file notification mode uses S3 event notifications published to an SNS topic, which fans out to an SQS queue that Auto Loader polls for new file events.
When `cloudFiles.inferColumnTypes` is set to false in Auto Loader schema inference, how are columns typed?
Answer: All inferred columns default to StringType
With `cloudFiles.inferColumnTypes=false`, Auto Loader infers the schema but types every column as StringType, which is the safest fallback for heterogeneous data.
Which statement BEST compares COPY INTO versus Auto Loader for data ingestion?
Answer: Auto Loader is better suited for large-scale continuous streaming; COPY INTO is designed for idempotent batch loads
Auto Loader excels at large-scale, continuously arriving data with Structured Streaming, while COPY INTO is an idempotent SQL command suited for scheduled batch ingestion into Delta tables.
Which of the following file formats is NOT natively supported as input by Auto Loader?
Answer: Microsoft Excel (.xlsx)
Auto Loader natively supports JSON, CSV, Parquet, Avro, ORC, text, and binary files, but does not support proprietary formats like Microsoft Excel (.xlsx).
What is the purpose of the `rescuedDataColumn` option in Auto Loader?
Answer: It stores rows or fields that could not be parsed due to schema mismatches in a dedicated column
`rescuedDataColumn` causes Auto Loader to place any data that doesn't match the schema (unexpected columns, type mismatches) into a named column as a JSON string rather than dropping or failing.
How does Auto Loader prevent duplicate file processing when a streaming job restarts after a failure?
Answer: It uses a checkpoint directory to track previously processed files, ensuring each file is processed exactly once
Auto Loader's checkpoint directory stores offsets and the complete list of processed files, so on restart it knows exactly which files have already been ingested and skips them.