โ† All CSI Flashcard Decks

Data Mapping & Transformation Flashcards

7 cards from real CSI practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Mapping & Transformation flashcards as text
  1. Which transformation pattern converts a single source record into multiple target records?

    Answer: One-to-many splitting

    One-to-many splitting (or explosion) breaks a single source record into multiple target records, such as expanding a line-item order into individual product records.

  2. In XSLT-based data transformation, what is the primary role of a stylesheet?

    Answer: Defining rules for transforming XML documents into another format

    An XSLT stylesheet defines templates and transformation rules that convert XML source documents into a target format such as HTML, another XML structure, or plain text.

  3. When performing a many-to-one aggregation in data mapping, which operation is most commonly applied?

    Answer: Combining multiple source records into a single target record using a function like SUM or COUNT

    Many-to-one aggregation consolidates multiple source records into a single target record, typically using aggregate functions such as SUM, AVG, MAX, or COUNT.

  4. A canonical data model in enterprise integration is best described as:

    Answer: A common, application-neutral data format used as an intermediary for message translation

    A canonical data model (CDM) is a vendor-neutral, agreed-upon intermediate format that decouples source and target systems, reducing the number of point-to-point transformations needed.

  5. Which data quality issue occurs when a field expected to hold a date contains a string like 'N/A'?

    Answer: Data type mismatch

    A data type mismatch occurs when a value stored in a field does not conform to the expected data type, such as text appearing in a date column.

  6. In ETL pipelines, what is the purpose of a surrogate key introduced during the transformation phase?

    Answer: To provide a system-generated, unique identifier independent of source system keys

    Surrogate keys are system-generated integers or GUIDs assigned in the data warehouse to uniquely identify records regardless of source system key formats or changes.

  7. Which approach is used to handle schema evolution in data transformation pipelines with minimal disruption?

    Answer: Using flexible schema registries and forward/backward-compatible serialization formats

    Schema registries (e.g., Confluent Schema Registry) combined with formats like Avro or Protobuf support forward and backward compatibility, allowing schemas to evolve without breaking existing consumers.