Spark SQL and DataFrames Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Spark SQL and DataFrames flashcards as text
Which Spark SQL function is used to perform an inner join between two DataFrames?
Answer: df1.join(df2, 'key')
df1.join(df2, 'key') performs a join between two DataFrames; the default join type is inner.
What is the purpose of the groupBy() function in Spark DataFrames?
Answer: Groups rows by specified columns for aggregation
groupBy() groups DataFrame rows by one or more columns and is typically followed by an aggregation function like agg(), count(), or sum().
Which function converts a DataFrame column to a different data type in Spark SQL?
Answer: df.col('x').cast()
cast() is used to convert a column to a specified data type, e.g., col('age').cast('integer').
What is Catalyst in Apache Spark?
Answer: The query optimizer for Spark SQL
Catalyst is Spark SQL's extensible query optimizer that transforms logical plans into optimized physical plans.
Which Spark SQL function removes duplicate rows from a DataFrame?
Answer: dropDuplicates()
dropDuplicates() removes duplicate rows, optionally considering only a subset of columns.
How do you write a Spark DataFrame to a Parquet file?
Answer: df.write.parquet('path')
df.write.parquet('path') saves the DataFrame to the specified path in Parquet format.