โ† All MS-DS Master of Data science Flashcard Decks

Big Data Technologies Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Big Data Technologies flashcards as text
  1. In the Lambda architecture for big data, what is the purpose of the 'speed layer'?

    Answer: To provide low-latency views of recent data that the batch layer has not yet processed

    The speed layer handles real-time data streams to fill the latency gap left by the batch layer's processing delay.

  2. Which Apache HBase concept ensures that writes are not lost if a RegionServer crashes before flushing to disk?

    Answer: Write-Ahead Log (WAL)

    HBase's Write-Ahead Log records every mutation before it is applied to the MemStore, allowing recovery after a crash.

  3. When using Spark Structured Streaming, what does 'output mode: complete' mean?

    Answer: The entire result table is written to the sink on every trigger

    In complete mode, the full aggregated result table is rewritten to the output sink every time a trigger fires.

  4. What is the primary function of Apache ZooKeeper in a distributed big data system?

    Answer: Providing distributed coordination, leader election, and configuration management

    ZooKeeper offers a centralized service for maintaining configuration, naming, synchronization, and group services in distributed systems.

  5. In Apache Spark, what is 'data skew' and what is a common mitigation technique?

    Answer: Data skew is uneven partition sizes; a common fix is salting keys to redistribute load

    Data skew occurs when some partitions are much larger than others, and salting adds a random prefix to keys to spread the load more evenly.

  6. Which technique does Apache Kafka use to guarantee message ordering within a topic?

    Answer: Ordering is guaranteed only within a single partition

    Kafka guarantees message order only within a partition; to maintain order for a key, all messages with that key must be routed to the same partition.

  7. What is the role of the 'NameNode' in HDFS?

    Answer: It manages the filesystem namespace and metadata, tracking where blocks are stored across DataNodes

    The NameNode holds all filesystem metadata (file names, directory structure, block locations) but does not store actual data.