โ† All Data Engineering Flashcard Decks

Cloud Data Storage Solutions Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Cloud Data Storage Solutions flashcards as text
  1. Which Azure service provides hierarchical namespace support optimized for big-data analytics?

    Answer: Azure Data Lake Storage Gen2

    ADLS Gen2 adds a hierarchical namespace on top of Blob Storage for analytics workloads.

  2. What is the purpose of partitioning data in cloud object storage for a query engine?

    Answer: To prune irrelevant data and reduce scan volume

    Partitioning lets query engines skip directories that don't match filter predicates, cutting scan cost.

  3. Which consistency model does Amazon S3 now provide for read-after-write on new objects?

    Answer: Strong read-after-write consistency

    S3 provides strong read-after-write consistency for PUTs of new objects.

  4. What is a key reason to enable object versioning in a cloud bucket?

    Answer: To protect against accidental overwrites and deletions

    Versioning retains previous object copies, allowing recovery from unintended changes.

  5. Which open table format adds ACID transactions and time travel on top of a data lake?

    Answer: Apache Iceberg

    Apache Iceberg (like Delta Lake and Hudi) brings ACID transactions and snapshots to lake storage.

  6. What does 'data egress cost' refer to in cloud storage pricing?

    Answer: Cost to transfer data out of the cloud provider's network

    Egress charges apply when data leaves the provider's network, often to the internet or another region.

  7. Which approach minimizes small-file problems in a cloud data lake?

    Answer: Compacting many small files into larger files

    Compaction merges many small files into fewer large files, improving read throughput and reducing overhead.