Data Serialization Formats Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Serialization Formats flashcards as text
What is the primary advantage of Apache Parquet over CSV for analytical workloads?
Answer: It stores data in columnar format enabling faster analytical queries
Parquet stores data in a columnar format, allowing analytical queries to read only the relevant columns and dramatically reducing I/O.
Which serialization format uses binary encoding and requires a schema to serialize and deserialize data?
Answer: Apache Avro
Apache Avro uses binary encoding and requires a schema (defined in JSON) to serialize and deserialize data, making it compact and efficient.
What type of internal storage organization does ORC (Optimized Row Columnar) use?
Answer: A hybrid of row stripes containing columnar data
ORC splits data into stripes (row groups) and within each stripe stores data in columnar format, combining the benefits of both row and columnar approaches.
Which serialization format is most appropriate when humans need to inspect and edit data directly without special tools?
Answer: JSON
JSON is a human-readable text format that can be opened and understood in any text editor, making it ideal for debugging and manual inspection.
What does 'schema evolution' mean in the context of data serialization?
Answer: The ability to modify a schema over time while maintaining compatibility with existing data
Schema evolution is the ability to modify a schema (e.g., add or remove fields) over time while still reading data written with older or newer schema versions.
Which format provides the best compression ratios for analytical data due to columnar storage and value encoding techniques?
Answer: Parquet
Parquet achieves excellent compression by storing similar data types together in columns, enabling encodings like dictionary encoding and run-length encoding that exploit data locality.
In Apache Avro, where is the writer's schema typically stored to make files self-describing?
Answer: In the file header of each Avro data file
Avro embeds the writer's schema in the header of each data file, ensuring the schema travels with the data and enabling self-describing files.