Skip to content

pudl_diff

Compare PUDL Parquet outputs between two dataset roots.

A "root" is a local or remote directory containing the Parquet outputs of a full PUDL ETL run, along with a datapackage descriptor of those outputs (e.g. $PUDL_OUTPUT/parquet or s3://pudl.catalyst.coop/nightly). This package lets callers load and compare tables between two such roots, and report on the differences.

To compare two datasets, use run_dataset_diff(), which returns a PudlDiffReport. It is a Pydantic model, so model_dump_json() and model_validate_json() write and read the JSON report. The pudl_diff command line tool is a thin wrapper around it.

The modules are layered, each importing only from those before it in this list:

  • logs, datapackage and defaults: standard library logging, the parts of a datapackage descriptor that the tool reads, and the few things the tool knows about PUDL (its nightly build, and where its metadata is), which are all optional so that the tool doesn't depend on the rest of PUDL.
  • base, formatting, dataset and performance: the base class of the report's models, formatting sizes and durations, access to a dataset and its tables, and sampling memory and CPU use.
  • schema, row_counts and rows: the three comparisons of a pair of tables, from cheapest to most expensive.
  • table and outputs: running all of them on one table, and writing the differing rows to Parquet files.
  • table_report and dataset_report: the serializable reports on one table and on a whole dataset.
  • terminal: rendering reports as text for a terminal.
  • runner: comparing many tables and building the report.