ADE Cloud & Big Data Technologies Flashcards
6 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 ADE Cloud & Big Data Technologies flashcards as text
What is 'data partitioning' in the context of cloud storage and big data?
Answer: Dividing data into logical segments to improve query performance and manageability
Data partitioning divides large datasets into logical segments (e.g., by date or region) so queries scan only relevant partitions, improving performance.
In Spark, what transformation is used to combine two DataFrames based on a common key?
Answer: join()
The join() transformation in Spark combines two DataFrames on a specified key, similar to SQL JOIN operations.
Which concept describes the ability of a cloud data system to handle growing data volumes by adding more nodes?
Answer: Horizontal scaling
Horizontal scaling (scale-out) adds more nodes to a distributed system, enabling cloud data platforms like Spark or BigQuery to handle petabyte-scale growth.
What does HDFS stand for in the Hadoop ecosystem?
Answer: Hadoop Distributed File System
HDFS (Hadoop Distributed File System) is Hadoop's primary storage system, designed to store very large files across commodity hardware with replication.
In a Lambda architecture, what is the role of the speed layer?
Answer: To provide low-latency real-time views of recent data
The speed layer in Lambda architecture processes real-time data streams to provide low-latency views that compensate for the batch layer's high latency.
Which file format supports ACID transactions natively in a data lake environment?
Answer: Delta Lake (Delta format)
Delta Lake adds ACID transaction support on top of Parquet files in a data lake, enabling reliable concurrent reads and writes.