Databricks Certified Associate Developer for Apache Spark — Questions and Answers
Question 1: Which method forces Spark to recompute an RDD and remove its cached data?
- rdd.unpersist() (Correct answer)
- rdd.evict()
- rdd.clear()
- rdd.release()
Correct answer: rdd.unpersist()
unpersist() removes the RDD from the cache, freeing the memory for other computations.
Question 2: Which format is recommended for reading structured data in Spark SQL when schema inference is needed?
- ORC
- CSV
- JSON (Correct answer)
- Parquet
Correct answer: JSON
JSON allows Spark to automatically infer the schema during read, though Parquet is preferred for performance in production.
Question 3: What is the role of the SparkContext in a Spark application?
- Serves as the entry point for Spark functionality and represents the connection to a Spark cluster (Correct answer)
- Manages SQL queries
- Provides the DataFrame API
- Schedules streaming jobs
Correct answer: Serves as the entry point for Spark functionality and represents the connection to a Spark cluster
SparkContext is the main entry point for Spark Core API, representing the connection to a Spark cluster and used to create RDDs.
Question 4: What is the purpose of the groupBy() function in Spark DataFrames?
- Partitions the DataFrame by specified columns
- Groups rows by specified columns for aggregation (Correct answer)
- Filters rows that belong to the same group
- Sorts the DataFrame by specified columns
Correct answer: Groups rows by specified columns for aggregation
groupBy() groups DataFrame rows by one or more columns and is typically followed by an aggregation function like agg(), count(), or sum().
Question 5: What does the repartition() function do to a Spark DataFrame?
- Increases or decreases the number of partitions by performing a full shuffle (Correct answer)
- Reduces partitions without shuffling
- Saves the DataFrame with custom partition keys
- Merges small files into larger ones
Correct answer: Increases or decreases the number of partitions by performing a full shuffle
repartition() performs a full shuffle to create the specified number of partitions, evenly distributing the data.
Question 6: What is predicate pushdown in Spark SQL?
- Applying filters after all joins have been completed
- Pushing aggregation predicates to the driver for central computation
- Moving filter conditions to earlier stages of the query plan to reduce data processed (Correct answer)
- Sending filter conditions to the cluster manager for partition pruning
Correct answer: Moving filter conditions to earlier stages of the query plan to reduce data processed
Predicate pushdown moves filter conditions as close to the data source as possible, reducing the amount of data read from disk.
Question 7: How does Spark handle executor failures during a job?
- Spark re-schedules the failed tasks on other available executors using RDD lineage (Correct answer)
- The job fails immediately when any executor is lost
- Spark pauses the job and waits for the executor to recover
- The driver takes over the computations from the failed executor
Correct answer: Spark re-schedules the failed tasks on other available executors using RDD lineage
When an executor fails, Spark reschedules its tasks on other executors and recomputes any lost RDD partitions using lineage.
Question 8: What is the purpose of the BlockManager in Apache Spark?
- Manages HDFS block replication for Spark's data
- Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk (Correct answer)
- Manages disk I/O operations for reading and writing data files
- Blocks network access between executors for security
Correct answer: Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk
BlockManager is Spark's distributed storage system that manages how data (RDD partitions, shuffle blocks, broadcasts) is stored across memory and disk.
Question 9: Which Spark MLlib algorithm is suitable for large-scale linear regression?
- DecisionTreeRegressor
- RandomForestRegressor
- LinearRegression with L-BFGS optimizer (Correct answer)
- GBTRegressor
Correct answer: LinearRegression with L-BFGS optimizer
Spark's LinearRegression uses L-BFGS or OLS (ordinary least squares) optimization and scales well to large distributed datasets.
Question 10: What does the `triplets` property of a GraphX Graph return?
- Three separate RDDs for source, edge, and destination
- A tuple of (VertexRDD, EdgeRDD, PartitionRDD)
- An array of three graphs representing different structural views
- An RDD of EdgeTriplet objects containing source vertex, edge, and destination vertex attributes (Correct answer)
Correct answer: An RDD of EdgeTriplet objects containing source vertex, edge, and destination vertex attributes
The `triplets` property returns an RDD[EdgeTriplet[VD, ED]], where each triplet combines the source vertex, the connecting edge, and the destination vertex with all their attributes.
Question 11: What is the purpose of the VectorAssembler in Spark MLlib?
- Scales feature values to a standard range
- Encodes categorical features as binary vectors
- Combines multiple feature columns into a single vector column (Correct answer)
- Reduces the dimensionality of features
Correct answer: Combines multiple feature columns into a single vector column
VectorAssembler merges multiple numeric columns into a single feature vector required by Spark ML algorithms.
Question 12: How do you write a Spark DataFrame to a Parquet file?
- df.save('path', 'parquet')
- df.to_parquet('path')
- df.export('path', format='parquet')
- df.write.parquet('path') (Correct answer)
Correct answer: df.write.parquet('path')
df.write.parquet('path') saves the DataFrame to the specified path in Parquet format.
Question 13: Which Spark configuration sets the amount of memory allocated per executor?
- spark.memory.fraction
- spark.driver.memory
- spark.executor.memory (Correct answer)
- spark.executor.cores
Correct answer: spark.executor.memory
spark.executor.memory sets the amount of memory (e.g., '4g') allocated to each Spark executor JVM process.
Question 14: Which Spark MLlib algorithm is used for collaborative filtering-based recommendations?
- ALS (Alternating Least Squares) (Correct answer)
- KMeans
- RandomForest
- LinearRegression
Correct answer: ALS (Alternating Least Squares)
ALS (Alternating Least Squares) is Spark's collaborative filtering algorithm for building recommendation systems.
Question 15: What is dynamic resource allocation in Spark?
- Automatically adjusting executor memory based on task requirements
- Distributing data across nodes based on access frequency
- Automatically partitioning data based on cluster size
- Scaling the number of executors up and down based on workload (Correct answer)
Correct answer: Scaling the number of executors up and down based on workload
Dynamic resource allocation allows Spark to add executors when tasks are queued and remove idle executors to free cluster resources.
Question 16: In Spark Structured Streaming, which sink writes results to an in-memory table for interactive querying?
- console
- foreach
- memory (Correct answer)
- file
Correct answer: memory
The memory sink stores the output in-memory as a named table, suitable for testing and interactive queries.
Question 17: RDDs are immutable and fault-tolerant.
- False
- True (Correct answer)
Correct answer: True
RDDs (Resilient Distributed Datasets) are indeed immutable, meaning once created, their contents cannot be changed. They are also fault-tolerant because they can be rebuilt from their lineage of transformations if a partition is lost, ensuring data integrity. These properties are fundamental to Spark's reliability and performance.
Question 18: What is the role of the DAGScheduler in Apache Spark?
- Converts the RDD DAG into stages and submits task sets to the TaskScheduler (Correct answer)
- Manages the lifecycle of Spark executors
- Monitors cluster health and restarts failed nodes
- Handles shuffle data transfer between executors
Correct answer: Converts the RDD DAG into stages and submits task sets to the TaskScheduler
The DAGScheduler transforms the logical DAG of RDD operations into physical stages and submits sets of tasks to the TaskScheduler for execution.
Question 19: Which method is used to run a SQL query on a registered temporary view in Spark SQL?
- spark.execute()
- spark.sql() (Correct answer)
- spark.query()
- spark.runSQL()
Correct answer: spark.sql()
spark.sql() executes a SQL query and returns the results as a DataFrame.
Question 20: Spark supports which cluster managers?
- All of the above (Correct answer)
- YARN
- MESOS
- Standalone Cluster Manager
Correct answer: All of the above
Apache Spark is designed to be flexible regarding cluster management. It supports running on various cluster managers, including its own Standalone cluster manager, Apache Mesos, and Hadoop YARN, allowing users to choose the best environment for their needs. This broad compatibility makes Spark adaptable to different infrastructure setups.
Question 21: What is a Spark Stage?
- A group of tasks that can be executed in parallel without a shuffle
- A phase of the DAG separated by wide transformations (shuffle boundaries) (Correct answer)
- A single task running on one partition
- A container for multiple Spark jobs
Correct answer: A phase of the DAG separated by wide transformations (shuffle boundaries)
A Stage is a set of tasks that can be computed without shuffling data; stage boundaries are determined by wide transformations.
Question 22: Which of the following is a valid trigger for Structured Streaming in Spark?
- ProcessingTime trigger (Correct answer)
- BatchSize trigger
- EventTime trigger
- RowCount trigger
Correct answer: ProcessingTime trigger
ProcessingTime trigger runs micro-batches at a specified interval, e.g., Trigger.ProcessingTime('10 seconds').
Question 23: What is dynamic partition pruning in Spark 3.x?
- Pruning null values from partitioned datasets
- Filtering partitions of a fact table at runtime using values from a dimension table join (Correct answer)
- Automatically reducing partition count after shuffles
- Removing empty partitions from the output
Correct answer: Filtering partitions of a fact table at runtime using values from a dimension table join
Dynamic partition pruning pushes the result of a dimension table filter into the fact table scan, skipping irrelevant partitions at runtime.
Question 24: What is GBTClassifier in Spark MLlib?
- A Gaussian Bayesian Tree Classifier
- A Gradient Boosted Trees Classifier for binary classification (Correct answer)
- A generalized batch training algorithm
- A graph-based binary tree classifier
Correct answer: A Gradient Boosted Trees Classifier for binary classification
GBTClassifier implements gradient boosted trees for binary classification, which iteratively trains decision trees to correct previous errors.
Question 25: What are Spark Executors?
- Processes on the driver node that compile the DAG into tasks
- Memory managers that handle executor JVM garbage collection
- JVM processes launched by the cluster manager that execute tasks and store RDD data (Correct answer)
- Worker nodes in the cluster that manage data replication
Correct answer: JVM processes launched by the cluster manager that execute tasks and store RDD data
Executors are JVM processes running on worker nodes that execute tasks assigned by the Driver and cache RDD partitions in memory.
Question 26: Which function converts a DataFrame column to a different data type in Spark SQL?
- df.col('x').cast() (Correct answer)
- df.col('x').asType()
- df.col('x').transform()
- df.col('x').convert()
Correct answer: df.col('x').cast()
cast() is used to convert a column to a specified data type, e.g., col('age').cast('integer').
Question 27: Which Spark SQL function removes duplicate rows from a DataFrame?
- deduplicate()
- unique()
- removeDups()
- dropDuplicates() (Correct answer)
Correct answer: dropDuplicates()
dropDuplicates() removes duplicate rows, optionally considering only a subset of columns.
Question 28: In GraphX's `aggregateMessages`, what is the role of the `sendMsg` function?
- To send final computation results back to the driver
- To broadcast vertex attributes to all partitions
- To define what messages are sent along each edge to neighboring vertices (Correct answer)
- To serialize edge data for network transfer
Correct answer: To define what messages are sent along each edge to neighboring vertices
The `sendMsg` function is called for each edge triplet and specifies what messages, if any, to send to the source and/or destination vertices for aggregation.
Question 29: Which of the following strategies helps avoid data skew in a join operation?
- Reducing the number of executors
- Using broadcast join for both sides
- Salting the join key with a random prefix (Correct answer)
- Increasing shuffle partitions
Correct answer: Salting the join key with a random prefix
Salting adds a random prefix to skewed keys, distributing hot keys across multiple partitions to avoid overloading a single task.
Question 30: What is the entry point for Spark SQL functionality in Spark 2.x and later?
- HiveContext
- SQLContext
- SparkContext
- SparkSession (Correct answer)
Correct answer: SparkSession
SparkSession is the unified entry point for Spark SQL, DataFrame, and Dataset APIs introduced in Spark 2.0.
Question 31: Which method waits for a streaming query to finish in Spark?
- query.block()
- query.wait()
- query.awaitTermination() (Correct answer)
- query.join()
Correct answer: query.awaitTermination()
awaitTermination() blocks the driver program until the streaming query is stopped or encounters an error.
Question 32: What is the purpose of the IndexToString transformer in Spark ML?
- Converts string labels to numeric indices
- Converts numeric predictions back to original string labels (Correct answer)
- Maps categorical features to binary vectors
- Converts dense vectors to sparse format
Correct answer: Converts numeric predictions back to original string labels
IndexToString reverses the effect of StringIndexer, converting numeric label predictions back to the original string labels.
Question 33: Which GraphX edge partitioning strategy assigns edges to partitions based solely on the source vertex ID?
- CanonicalRandomVertexCut
- RandomVertexCut
- EdgePartition2D
- EdgePartition1D (Correct answer)
Correct answer: EdgePartition1D
EdgePartition1D assigns edges to partitions by hashing only the source vertex ID, colocating all outgoing edges from the same vertex in one partition.
Question 34: Which Spark MLlib algorithm performs dimensionality reduction?
- KMeans
- GBTClassifier
- PCA (Principal Component Analysis) (Correct answer)
- Naive Bayes
Correct answer: PCA (Principal Component Analysis)
PCA (Principal Component Analysis) reduces the dimensionality of feature vectors by projecting them onto principal components.
Question 35: Which storage level in Spark caches RDD data in memory as deserialized Java objects?
- MEMORY_AND_DISK_SER
- DISK_ONLY
- MEMORY_ONLY (Correct answer)
- MEMORY_ONLY_SER
Correct answer: MEMORY_ONLY
MEMORY_ONLY stores RDD data in memory as deserialized Java objects, offering the fastest access but highest memory usage.
Question 36: What does deploy mode 'cluster' mean in Spark?
- The application is replicated across multiple driver nodes for high availability
- The Driver process runs on one of the worker nodes managed by the cluster manager (Correct answer)
- The Spark application runs on every node in the cluster
- The cluster manager selects the optimal executor placement
Correct answer: The Driver process runs on one of the worker nodes managed by the cluster manager
In cluster deploy mode, the Driver runs on one of the worker nodes in the cluster, which is preferred for production jobs.
Question 37: What is the purpose of the `subgraph` operation in GraphX?
- To split a graph into two equal halves for parallel processing
- To create a coarser hierarchical summary of the graph
- To merge two separate graphs into one larger graph
- To extract a subset of vertices and edges that satisfy a given predicate (Correct answer)
Correct answer: To extract a subset of vertices and edges that satisfy a given predicate
The `subgraph` operation filters vertices and edges using user-supplied predicates, returning a new graph that contains only the elements satisfying both conditions.
Question 38: What is the role of offsets in Spark Structured Streaming with Kafka?
- They represent the number of records processed per second
- They define the batch size for each micro-batch
- They configure the Kafka partition assignment strategy
- They track the position in the Kafka topic for fault-tolerant exactly-once processing (Correct answer)
Correct answer: They track the position in the Kafka topic for fault-tolerant exactly-once processing
Offsets track the last-read position in Kafka topics, stored in checkpoints, enabling the stream to resume from where it stopped.
Question 39: What is the purpose of the external shuffle service in Spark?
- Compresses shuffle data to reduce network bandwidth usage
- Handles shuffle operations faster than in-executor processing
- Maintains shuffle files on worker nodes so executors can be released during dynamic allocation (Correct answer)
- Streams shuffle data directly between executors bypassing the driver
Correct answer: Maintains shuffle files on worker nodes so executors can be released during dynamic allocation
The external shuffle service runs on worker nodes and holds shuffle files, allowing executors to be released while preserving shuffle data.
Question 40: Which feature transformer in Spark ML converts text documents to TF-IDF feature vectors?
- HashingTF + IDF (Correct answer)
- Word2Vec
- NGram
- BertEncoder
Correct answer: HashingTF + IDF
HashingTF converts text tokens to term frequency vectors, and IDF scales by inverse document frequency to produce TF-IDF features.
Question 41: What does the coalesce() transformation do in Spark?
- Caches the RDD to memory
- Reduces the number of partitions with minimal data shuffling (Correct answer)
- Increases the number of partitions
- Merges two RDDs into one
Correct answer: Reduces the number of partitions with minimal data shuffling
coalesce() reduces the number of partitions by merging them, avoiding a full shuffle when decreasing partition count.
Question 42: In which deploy mode does the Spark driver run on the machine where spark-submit is executed?
- Standalone mode
- Cluster mode
- Client mode (Correct answer)
- Local mode
Correct answer: Client mode
In client mode, the Driver runs on the machine that submitted the application, which is useful for interactive development but not recommended for production.
Question 43: How does Spark achieve fault tolerance with RDDs?
- By replicating data across multiple nodes
- By writing all intermediate data to disk
- By recomputing lost partitions using lineage information (Correct answer)
- By maintaining checkpoints after every transformation
Correct answer: By recomputing lost partitions using lineage information
Spark achieves fault tolerance by recomputing lost RDD partitions using the recorded lineage of transformations.
Question 44: Which source reads data from Apache Kafka in Spark Structured Streaming?
- spark.kafka.read()
- spark.readStream.format('kafka') (Correct answer)
- spark.readStream.format('kafkaSource')
- spark.stream('kafka')
Correct answer: spark.readStream.format('kafka')
spark.readStream.format('kafka') reads streaming data from Kafka topics using the Kafka connector for Structured Streaming.
Question 45: What is the function of withColumn() in Spark DataFrames?
- Adds a new column or replaces an existing column with the given column expression (Correct answer)
- Selects only the specified column
- Renames a column
- Removes a column from the DataFrame
Correct answer: Adds a new column or replaces an existing column with the given column expression
withColumn() returns a new DataFrame by adding or replacing a column using the provided expression.
Databricks Certified Associate Developer for Apache Spark
The Databricks Certified Associate Developer for Apache Spark certification validates proficiency in using Apache Spark and the Databricks Lakehouse Platform to complete introductory-level data engineering tasks, covering Spark architecture, SQL, DataFrames, performance tuning, and advanced processing features.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds