Azure Databricks Flashcards
7 cards from real DP-900 practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Azure Databricks flashcards as text
What is the Azure Databricks workspace?
Answer: An environment that organizes notebooks, clusters, jobs, and data into a unified interface
The Azure Databricks workspace is a unified environment that organizes assets such as notebooks, clusters, jobs, libraries, and data, enabling team collaboration.
Which feature of Delta Lake allows users to query data as it existed at a previous point in time?
Answer: Time travel
Delta Lake's time travel feature uses a transaction log to let users query historical versions of data using a timestamp or version number.
What is MLflow in the context of Azure Databricks?
Answer: An open-source platform for managing the machine learning lifecycle including experiments and model registry
MLflow is an open-source platform integrated into Azure Databricks that helps manage the ML lifecycle, including experiment tracking, model versioning, and deployment.
How are Azure Databricks clusters billed?
Answer: Based on Databricks Units (DBUs) consumed during cluster operation
Azure Databricks uses a consumption-based model where you are billed for Databricks Units (DBUs), which represent processing capability per hour, combined with underlying VM costs.
What is a Databricks Job?
Answer: A scheduled or triggered execution of a notebook or JAR on a cluster
A Databricks Job is a way to schedule or trigger the automated execution of notebooks, Python scripts, or JAR files on a cluster for production workloads.
Which two types of clusters does Azure Databricks offer?
Answer: All-purpose clusters and job clusters
Azure Databricks offers all-purpose clusters for interactive development and collaboration, and job clusters that are created for a single job and terminated after the job completes.
What file format does Delta Lake use to physically store data?
Answer: Parquet
Delta Lake stores data in Parquet file format, which is a columnar storage format optimized for analytics workloads, combined with a transaction log for ACID compliance.