ADE Cloud & Big Data Technologies Flashcards
6 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 ADE Cloud & Big Data Technologies flashcards as text
What does the term 'data lakehouse' refer to?
Answer: An architecture combining the flexibility of data lakes with the management features of data warehouses
A data lakehouse merges the low-cost storage of data lakes with ACID transactions, schema enforcement, and BI support from data warehouses.
Which Apache technology is used for distributed message streaming and serves as a backbone for real-time data pipelines?
Answer: Apache Kafka
Apache Kafka is a distributed event streaming platform used to build real-time data pipelines and streaming applications at scale.
In cloud-based big data architectures, what is 'schema-on-read'?
Answer: Defining the schema when data is read rather than when it is stored
Schema-on-read means data is stored in raw form and the schema is applied at query time, providing flexibility for diverse data types.
Which Azure service is used to orchestrate large-scale data movement and transformation pipelines?
Answer: Azure Data Factory
Azure Data Factory is a cloud-based ETL and data integration service for creating data-driven workflows to orchestrate data movement and transformation.
What is the primary advantage of using columnar storage over row-based storage for analytical workloads?
Answer: Improved read performance for aggregate queries on specific columns
Columnar storage allows analytical queries to read only the columns needed, drastically reducing I/O and improving aggregate query performance.
Which AWS service provides a fully managed ETL service that automatically generates ETL code?
Answer: AWS Glue
AWS Glue is a serverless ETL service that discovers data schemas via crawlers and auto-generates PySpark or Scala ETL scripts.