โ† All Data Engineering Flashcard Decks

Data Serialization Formats Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Serialization Formats flashcards as text
  1. Which encoding technique does Parquet use for columns with few distinct values (low cardinality)?

    Answer: Dictionary encoding

    Dictionary encoding replaces repeated values with integer references to a small dictionary, which is highly effective for low-cardinality columns where the same values appear many times.

  2. What is the primary benefit of reading columnar formats like Parquet in analytical query engines?

    Answer: Only relevant columns are read from disk, reducing I/O significantly

    Columnar storage allows query engines to read only the columns they need and skip all others on disk, dramatically reducing I/O and speeding up analytical queries.

  3. Which serialization format is most widely used for data interchange with REST APIs?

    Answer: JSON

    JSON is the de facto standard for REST API data interchange due to its human-readable nature, wide language support, and native compatibility with web technologies.

  4. What is Protocol Buffers (protobuf) primarily designed for?

    Answer: Efficient binary serialization for structured data across languages

    Protocol Buffers is Google's binary serialization format designed for efficient, language-neutral structured data serialization with small payload size and fast parsing.

  5. Which Parquet feature allows query engines to skip entire row groups based on filter conditions without reading the data?

    Answer: Row group filtering using min/max statistics

    Parquet stores min/max statistics for each row group and column chunk, allowing query engines to skip entire row groups when filter conditions cannot match any row in that group.

  6. What is the key advantage of Avro for Apache Kafka message serialization compared to JSON?

    Answer: Its compact binary format and schema evolution support via a schema registry

    Avro's compact binary encoding reduces Kafka message size for higher throughput, and its schema evolution support with a schema registry allows producers and consumers to evolve independently.

  7. What is the difference between backward compatibility and forward compatibility in schema evolution?

    Answer: Backward: new schema reads old data; Forward: old schema reads new data

    Backward compatibility means a newer schema version can read data written by an older schema, while forward compatibility means an older schema can read data written by a newer schema.