โ† All Data Engineering Flashcard Decks

Fundamentals Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Fundamentals flashcards as text
  1. Three different types of data models are listed below, EXCEPT:

    Answer: Schema Data Model

    The three primary types of data models are Conceptual, Logical, and Physical. A Conceptual Data Model defines high-level business concepts, a Logical Data Model details entities and relationships without specific database technology, and a Physical Data Model describes how data is stored in a specific database system. 'Schema Data Model' is not a standard, distinct type in this widely accepted classification.

  2. The process of normalizing a data model is one of the steps in data modeling, particularly in...

    Answer: Logical Data Model

    Normalization is a crucial process primarily applied during the creation of a Logical Data Model. It involves organizing the columns and tables of a relational database to minimize data redundancy and improve data integrity. This step ensures the logical structure of the database is efficient, consistent, and adheres to best practices before moving to physical implementation.

  3. We need data engineers for a number of reasons, EXCEPT:

    Answer: Various Capabilities

    Data engineers are essential due to the complexities arising from various data formats (structured, unstructured), diverse data sources (databases, APIs, streaming), and a multitude of technologies (cloud platforms, big data tools). 'Various Capabilities' refers to the skills of the engineers themselves, not a reason for the *need* for engineers in the face of data challenges, which are driven by the data's characteristics and sources.

  4. Which of the following is NOT a part of the ETL / ELT process:

    Answer: Export

    ETL stands for Extract, Transform, Load, which are the three core phases of moving data from source systems to a data warehouse or data lake. 'Export' is not a standard component of the ETL/ELT acronym. While data might be exported at various stages in a broader data pipeline, it is not part of the fundamental definition of the ETL process itself.

  5. The following are components of the Data Pipeline in which Data Engineers play a significant role in:

    Answer: Collect and Prepare Data

    Data engineers play a significant role in the initial stages of the data pipeline, which involve collecting raw data from various sources and preparing it for analysis. This preparation includes cleaning, transforming, and structuring the data to make it usable for data scientists and analysts. Their expertise ensures data quality and accessibility for subsequent analytical processes.

  6. Which of the following is NOT a common source of big data?

    Answer: Official Statistics Surveys

    Big data typically refers to datasets characterized by high volume, velocity, and variety, often requiring specialized tools for processing. Common sources include sensors (IoT data), social media feeds, and web logs (like Google Trends data). Official statistics surveys, while producing large datasets, are generally structured, planned, and collected through traditional methods, and are not typically categorized as 'big data' in the context of its defining challenges.

  7. Which of the following commands belongs to the data manipulation language?

    Answer: Select, Insert, Delete

    The Data Manipulation Language (DML) in SQL includes commands used to manage data within database objects. `SELECT` retrieves data, `INSERT` adds new data, `UPDATE` modifies existing data, and `DELETE` removes data. These commands are specifically designed for interacting with and manipulating the data stored in the database tables.