← All Data and Analytics Flashcard Decks

Big Data and Cloud Analytics Flashcards

6 cards from real Data and Analytics practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Big Data and Cloud Analytics flashcards as text
  1. What is Google BigQuery's primary architectural advantage?

    Answer: Serverless, scalable analytics separating storage from compute

    BigQuery's serverless architecture separates storage from compute, allowing independent scaling and eliminating infrastructure management.

  2. What is Apache Kafka primarily used for?

    Answer: Real-time data streaming and event-driven pipelines

    Apache Kafka is a distributed event streaming platform used for high-throughput, real-time data pipelines and stream processing.

  3. What is the purpose of a CDN (Content Delivery Network) in analytics applications?

    Answer: To reduce latency by caching content closer to users geographically

    A CDN distributes cached content across geographically dispersed servers to reduce latency and improve load times for end users.

  4. What does 'schema-on-read' mean in the context of data lakes?

    Answer: Schema is applied when data is queried, not when it's stored

    Schema-on-read means raw data is stored without a fixed structure, and the schema is applied dynamically when the data is queried.

  5. What is Snowflake's multi-cluster, shared data architecture?

    Answer: Separate storage and compute layers allowing multiple independent query engines

    Snowflake separates storage from compute, allowing multiple virtual warehouses to independently access the same data without contention.

  6. What is the purpose of data partitioning in big data systems?

    Answer: To divide large datasets into smaller chunks for faster parallel processing

    Partitioning divides data into smaller, manageable chunks that can be processed in parallel across multiple nodes, improving query performance.