โ† All ACP Flashcard Decks

ACP Data Engineering & Workflow Automation Flashcards

6 cards from real ACP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 ACP Data Engineering & Workflow Automation flashcards as text
  1. Which Python library provides the `Parquet` file format support for high-performance columnar storage of DataFrames?

    Answer: pyarrow

    `pyarrow` (and `fastparquet`) enables pandas to read/write Parquet files via `df.to_parquet()` and `pd.read_parquet()`, offering efficient columnar compression for large datasets.

  2. Which conda command exports a complete list of all packages in the current environment to a file for reproducibility?

    Answer: conda env export > environment.yml

    `conda env export > environment.yml` captures all packages, versions, and channels in the active environment into a YAML file that can recreate the environment elsewhere.

  3. In a data engineering context, what is the purpose of using Dask instead of pandas for large dataset processing in Python?

    Answer: Dask enables parallel and out-of-core computation on datasets larger than RAM

    Dask partitions large datasets into chunks and processes them in parallel across cores or a cluster, handling data that exceeds available memory with a pandas-compatible API.

  4. Which method in pandas handles missing values by filling them with a specified value or strategy (e.g., forward fill)?

    Answer: df.fillna()

    `df.fillna(value)` replaces NaN values with a constant or uses methods like `ffill` (forward fill) and `bfill` (backward fill) to propagate adjacent valid values.

  5. When automating a data pipeline on a schedule using cron or a task scheduler, which Python standard library module provides programmatic access to run shell commands and subprocesses?

    Answer: subprocess

    The `subprocess` module provides `subprocess.run()` and `Popen` for launching external processes, capturing output, and handling errors within automated Python pipeline scripts.

  6. In pandas, which method converts a column's data type (e.g., from object/string to numeric or datetime)?

    Answer: df['col'].astype()

    `df['col'].astype(dtype)` casts a Series to the specified dtype (e.g., `int64`, `float32`, `str`), and `pd.to_datetime()` / `pd.to_numeric()` handle specialized conversions.