Big Data Architect Flashcards
7 cards from real IBM Certification practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data Architect flashcards as text
Which IBM data architecture pattern uses a 'data vault' modeling approach to improve auditability and adaptability in enterprise data warehouses?
Answer: Data Vault 2.0 with hubs, links, and satellites
Data Vault 2.0 organizes data into hubs (business keys), links (relationships), and satellites (context/attributes), enabling auditability and schema flexibility.
An architect designs an IBM data pipeline where Apache Spark jobs process HDFS data. What is the performance advantage of using Spark's in-memory processing over MapReduce?
Answer: Spark keeps intermediate data in memory across stages, reducing disk I/O significantly
Spark caches intermediate results in memory across transformation stages, dramatically reducing the disk I/O that MapReduce incurs between each map and reduce step.
In IBM Cloud Object Storage, what erasure coding configuration provides the best balance between storage efficiency and fault tolerance?
Answer: 10+4 erasure coding (10 data slices, 4 parity slices)
10+4 erasure coding stores data across 14 slices with 4 parity slices, tolerating 4 simultaneous failures while using less storage than 3x replication.
When integrating IBM MQ with a big data pipeline, what problem does persistent message delivery primarily solve?
Answer: Guaranteeing message delivery even if downstream consumers are temporarily unavailable
IBM MQ's persistent delivery ensures messages survive broker restarts and consumer outages, providing at-least-once delivery guarantees for critical data pipelines.
A Big Data Architect must enforce column-level security on an IBM Db2 Warehouse table containing PII. Which mechanism is most appropriate?
Answer: Row permissions using IBM row and column access control (RCAC)
IBM Db2 Row and Column Access Control (RCAC) allows fine-grained policies that restrict which rows and columns specific users or roles can access.
Which IBM tool provides automated discovery of sensitive data (PII, PHI, PCI) across heterogeneous big data sources to support compliance requirements?
Answer: IBM Watson Knowledge Catalog with data discovery rules
IBM Watson Knowledge Catalog includes automated data discovery with configurable rules that identify and classify sensitive data across connected sources.
In designing an IBM big data solution for a financial institution, why is data lineage tracking considered a critical architectural requirement?
Answer: It enables auditors to trace how any data point was derived and transformed end-to-end
Data lineage provides a complete audit trail of data origin, transformations, and movement, which is essential for regulatory compliance and impact analysis.