Data Mapping & Transformation Flashcards
7 cards from real CSI practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Mapping & Transformation flashcards as text
What is the role of an 'identity transformation' in a data pipeline?
Answer: It passes data through unchanged, serving as a passthrough or baseline for testing other transforms
An identity transformation passes source data to the target unchanged; it is useful for testing pipeline infrastructure, verifying connectivity, or as a placeholder before custom logic is added.
When mapping date fields between systems using different formats (e.g., MM/DD/YYYY vs. ISO 8601), which risk must be explicitly addressed?
Answer: Ambiguous date parsing leading to incorrect values, especially near month boundaries
Mismatched date formats can cause silent data errors—e.g., 04/05/2024 interpreted as April 5 in one system and May 4 in another—so explicit format parsing must be enforced.
Which ETL anti-pattern occurs when transformation logic is embedded directly in SQL stored procedures inside the target database rather than in a dedicated transformation layer?
Answer: Logic leakage into the database layer
Embedding transformation logic in stored procedures tightly couples business rules to the database engine, making the pipeline harder to test, version, and migrate to new platforms.
In data mapping documentation, what does a 'source-to-target mapping specification' typically include?
Answer: Source field name, target field name, data type, transformation logic, and business rules for each mapped field
A source-to-target mapping specification (or mapping document) catalogs every field relationship including source/target names, types, transformation rules, and validation logic.
Which technique splits a single field containing delimited values (e.g., 'red;blue;green') into separate records or columns?
Answer: String tokenization and unpivoting
String tokenization splits a delimited value by its separator, and unpivoting (or normalization) then creates individual rows or columns for each token.
What is 'data enrichment' in the context of data transformation pipelines?
Answer: Augmenting source records with additional attributes from external or reference datasets
Data enrichment enhances incoming records by appending additional information from reference databases, APIs, or lookups—such as adding geolocation data to an address field.
A 'pivot transformation' in data integration is used to:
Answer: Rotate data from row-based to column-based layout (or vice versa) based on values in a key column
A pivot transformation converts row-based data into columns (or unpivot does the reverse), restructuring the dataset layout to match the target schema requirements.