Cloud Data Storage Solutions Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Cloud Data Storage Solutions flashcards as text
What distinguishes a data lakehouse from a traditional data lake?
Answer: It combines lake-scale storage with warehouse-style management and ACID tables
A lakehouse layers warehouse features like ACID transactions and governance onto low-cost lake storage.
Which storage choice best supports globally distributed, low-latency reads of a content catalog?
Answer: Object storage fronted by a CDN
A CDN caches objects near users worldwide, delivering low-latency reads from object storage origins.
What is the purpose of a manifest or metadata layer in formats like Delta Lake?
Answer: To track which files belong to a table version for consistent reads
The metadata/transaction log records the set of files per snapshot, enabling ACID reads and time travel.
Which scenario most justifies using a cold storage tier?
Answer: Compliance records that must be retained but are rarely accessed
Cold tiers minimize storage cost for infrequently accessed data like long-term compliance archives.
What does 'storage tiering' aim to optimize in a cloud environment?
Answer: The balance between access cost, retrieval latency, and storage price
Tiering matches data to the storage class whose cost and latency fit its access pattern.
Which technique helps a query engine skip files without opening them in a cloud data lake?
Answer: Min/max column statistics and partition pruning
File-level min/max statistics let engines skip files that cannot contain matching rows.
What is a key consideration when choosing block size or file size for analytical cloud storage?
Answer: Balancing parallelism against per-file overhead, often hundreds of MB per file
Moderately large files (e.g., 128 MBโ1 GB) balance read parallelism with metadata and open overhead.