โ† All IBM Certification Flashcard Decks

Big Data Architect Flashcards

7 cards from real IBM Certification practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Big Data Architect flashcards as text
  1. In IBM Cloud Pak for Data, what is the primary purpose of the DataStage component?

    Answer: ETL and data integration pipelines

    IBM DataStage is an ETL tool within Cloud Pak for Data used to design, develop, and run data integration jobs.

  2. Which consistency model does Apache Cassandra use, and how does it align with IBM big data architecture recommendations for high availability?

    Answer: Eventual consistency with tunable read/write levels

    Cassandra uses eventual consistency by default but allows tunable consistency levels per operation, enabling architects to balance availability and consistency.

  3. When implementing data tiering in an IBM big data solution, which storage tier is most appropriate for infrequently accessed cold data?

    Answer: Object storage such as IBM Cloud Object Storage

    Object storage like IBM Cloud Object Storage is cost-effective and highly durable, making it ideal for cold or archival data tiers.

  4. In IBM Db2 Big SQL, what technique allows queries to span both HDFS-resident data and traditional Db2 relational tables in a single statement?

    Answer: Federated query with external table mappings

    Big SQL supports federated queries using external table definitions that map HDFS data sources alongside local Db2 tables.

  5. An architect is designing an IBM streaming solution where out-of-order events are common. Which windowing strategy best handles late-arriving data?

    Answer: Watermark-based event-time windows

    Watermark-based event-time windowing tracks the progress of event time and allows configurable tolerance for late-arriving data.

  6. Which IBM platform component provides automated data quality profiling and cleansing rules as part of a governed data pipeline?

    Answer: IBM InfoSphere Information Analyzer

    IBM InfoSphere Information Analyzer profiles data assets to detect quality issues, anomalies, and patterns across enterprise data sources.

  7. In a multi-zone IBM Cloud big data deployment, what is the main architectural goal of distributing HDFS DataNodes across availability zones?

    Answer: Ensuring rack-aware replication survives zone failure

    Distributing DataNodes across availability zones ensures that HDFS rack-aware replication policies protect data even if an entire zone fails.