Distributed Data Processing Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Distributed Data Processing flashcards as text
What is the benefit of partition pruning in a distributed query engine?
Answer: It skips reading partitions irrelevant to the filter
Partition pruning avoids scanning data that cannot match the query predicate.
In the lambda architecture, what is the speed layer responsible for?
Answer: Low-latency views of recent data
The speed layer serves real-time results while the batch layer recomputes accurate views.
Which factor most directly limits the scalability of a global ORDER BY in distributed SQL?
Answer: All data must funnel through a single sort/merge step
Total ordering forces a global merge that becomes a bottleneck at scale.
What does eventual consistency guarantee?
Answer: Replicas converge to the same value if updates stop
Given no new writes, all replicas eventually reflect the same data.
Why are columnar formats often paired with predicate pushdown?
Answer: The engine can skip blocks using column statistics
Min/max statistics per column block let the engine skip non-matching data.
In a consistent hashing ring, what happens when one node is removed?
Answer: Only its keys remap to the next node
Consistent hashing limits remapping to the departing node's keys, minimizing disruption.
What is the main reason to colocate compute with data in a cluster?
Answer: Reduce network transfer by processing data locally
Data locality minimizes network I/O by running tasks where the data already resides.