Big Data Technologies Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data Technologies flashcards as text
In Apache Spark, what is the 'Catalyst optimizer' responsible for?
Answer: Transforming logical query plans into optimized physical execution plans for DataFrames and Datasets
Catalyst is Spark SQL's rule-based and cost-based query optimizer that rewrites logical plans into efficient physical plans.
What distinguishes an 'append-only' log in Apache Kafka from a traditional message queue?
Answer: Kafka retains messages for a configurable retention period, allowing multiple consumers to replay events independently
Kafka's immutable log persists records beyond consumption, enabling multiple independent consumers and event replay without impacting each other.
Which Hadoop ecosystem tool is best suited for running SQL-like queries directly on data stored in HDFS or S3 without moving it to a separate database?
Answer: Apache Hive
Apache Hive provides HiveQL, a SQL dialect that translates queries into MapReduce, Tez, or Spark jobs that run directly over data in HDFS or compatible object stores.
In the context of big data pipelines, what is 'backpressure' and why does it matter?
Answer: A flow-control mechanism where a downstream component signals a slower ingestion rate to prevent being overwhelmed
Backpressure propagates downstream capacity constraints upstream, preventing faster producers from overwhelming slower consumers and causing out-of-memory failures.
What is the purpose of Delta Lake's transaction log ('_delta_log' directory)?
Answer: It records every change to the table as a JSON commit entry, enabling ACID transactions and time-travel queries
Delta Lake's transaction log is an ordered list of atomic commits that provides ACID properties, optimistic concurrency control, and the ability to query previous table versions.
Which partitioning strategy in Apache Kafka ensures that all messages with the same key always go to the same partition?
Answer: Hash-based key partitioning
Kafka's default partitioner computes a hash of the message key modulo the number of partitions, routing identical keys to the same partition deterministically.
In a distributed system using Apache Spark on a cluster, what is a 'shuffle' operation and why can it be expensive?
Answer: A shuffle redistributes data across partitions based on a key, requiring network data transfer between all executors
Shuffles involve serializing, transferring, and deserializing data across the network to group records by key, making them the most common performance bottleneck in Spark jobs.