ACP Data Engineering & Workflow Automation Flashcards
6 cards from real ACP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 ACP Data Engineering & Workflow Automation flashcards as text
Which Python library provides the `Parquet` file format support for high-performance columnar storage of DataFrames?
Answer: pyarrow
`pyarrow` (and `fastparquet`) enables pandas to read/write Parquet files via `df.to_parquet()` and `pd.read_parquet()`, offering efficient columnar compression for large datasets.
Which conda command exports a complete list of all packages in the current environment to a file for reproducibility?
Answer: conda env export > environment.yml
`conda env export > environment.yml` captures all packages, versions, and channels in the active environment into a YAML file that can recreate the environment elsewhere.
In a data engineering context, what is the purpose of using Dask instead of pandas for large dataset processing in Python?
Answer: Dask enables parallel and out-of-core computation on datasets larger than RAM
Dask partitions large datasets into chunks and processes them in parallel across cores or a cluster, handling data that exceeds available memory with a pandas-compatible API.
Which method in pandas handles missing values by filling them with a specified value or strategy (e.g., forward fill)?
Answer: df.fillna()
`df.fillna(value)` replaces NaN values with a constant or uses methods like `ffill` (forward fill) and `bfill` (backward fill) to propagate adjacent valid values.
When automating a data pipeline on a schedule using cron or a task scheduler, which Python standard library module provides programmatic access to run shell commands and subprocesses?
Answer: subprocess
The `subprocess` module provides `subprocess.run()` and `Popen` for launching external processes, capturing output, and handling errors within automated Python pipeline scripts.
In pandas, which method converts a column's data type (e.g., from object/string to numeric or datetime)?
Answer: df['col'].astype()
`df['col'].astype(dtype)` casts a Series to the specified dtype (e.g., `int64`, `float32`, `str`), and `pd.to_datetime()` / `pd.to_numeric()` handle specialized conversions.