PUDL Diff Changelog¶
v0.1.2 (unreleased)¶
v0.1.1 (2026-10-07)¶
2026-10-07¶
Command line¶
- The Python API now compares tables of any size by default, since the work on memory use means that no table needs a cutoff to be compared safely.
The command line still defaults
--max-compare-rowsto 100 million, so that a default comparison doesn't spend minutes on the very largest tables; pass a larger number, or 0 for no limit, to include them.DiffOptions.max_compare_rowscan now beNone(its default) for no limit, so the report schema version is now 1.1.0.162ac26 --from-reportnow accepts remote locations, such asgs://ors3://URLs, as well as local paths, usingUPathas the datasets already do. This allows looking at the report of a build without downloading it first.a656faa
v0.1.0 (2026-10-07)¶
pudl_diff was written in a few weeks of local development as a side-project before it was published on GitHub.
This means there are no issues or pull requests through which to trace its early history, and the changelog below stands in for them.
See also the implementation plan, which was the project's working plan in development.
2026-10-07¶
Documentation and housekeeping¶
- The README now has the status badges, installation instructions for pixi, uv, pip and conda, and links to the related PUDL and Catalyst projects, ready for the repository to be published.
2d4c25c - The implementation notes moved into
notes/, and are checked and formatted like the rest of the Markdown. The formatter now numbers lists consecutively, because the notes refer to their steps by number.92e42c5 - The release notes became this changelog.
092bb02 - The pixi lockfile was refreshed.
88ec2a3
API documentation¶
- The API reference is now generated by Zensical itself, which has gained native versions of the
mkdocs-api-autonavandmkdocs-autoapiplugins since the project was set up. A page for each module of the package, and its place in the navigation, come from the package's own source, so a new module no longer has to be added to a list by hand, anddocs/reference.mdis gone. The docstrings are still rendered by mkdocstrings, which Zensical doesn't replace, but the cross-references between pages are resolved by Zensical's own autorefs, which also publishes anobjects.invthat other projects' documentation can link to.67cd1dd - The changelog now links the classes, functions, constants and modules it mentions to the API reference, and the documentation build is strict, so that a reference that doesn't resolve fails it instead of silently becoming plain text.
67cd1dd - The API reference's navigation lists plain module names. The default "mod" badge in front of each was not spaced from its title by the theme, and its markup also leaked into the labels of the previous and next page links.
741850b - The implementation plan is part of the documentation, as a historical document, instead of a file in the repository.
67cd1dd - The spell checker ignores the commit and compare URLs and abbreviated hashes of the changelog, which contain fragments that look like misspelled words, but still checks the rest of it.
67cd1dd
Windows¶
- The tests failed on Windows, and one of the causes was a bug in the tool: it read and wrote datapackage descriptors and reports in the platform's default encoding, which on Windows is not UTF-8.
PUDL's descriptors have non-ASCII characters in their descriptions, and reports are archived and shown again elsewhere, so a report written on Linux could not be read on Windows.
All of the text the tool reads and writes is now UTF-8, and the generated schema files are written with the same line endings on every platform.
7f43879 - The other failures were in the tests, which compared paths as strings with the wrong separators, and expected
HOMEto set the home directory, which Windows takes fromUSERPROFILE.7f43879
2026-10-06¶
Comparing the largest tables¶
- The biggest table in PUDL,
core_epacems__hourly_emissions, has a billion rows, and comparing it took about 75 GB of memory at its peak. That was more than most developers' machines have. The cause was the step that reduces each table to one narrow row per key hash, which holds a group for every row on both sides at once. - Tables of more than 100 million rows (
MAX_ROWS_PER_PARTITION) are now reduced and joined in several passes, each keeping only the key hashes in one slice of the hash space. Rows with equal keys always land in the same slice, so the slices can be compared independently and their results combined. Peak memory is then about 8 GB per 100 million rows, however large the table is. - Partitioning on the hash of the key needs nothing that is specific to a table, and works with or without a primary key. The price is time, since every pass reads and hashes the whole table again, because a filter on a hash can't be pushed down into the Parquet reader. The billion-row table now takes 9.7 GB and about two minutes, instead of 75 GB and a minute.
- Tables of 100 million rows or fewer are compared in a single pass, as before.
Tables this large are still left out of a default comparison, since they exceed the default
--max-compare-rows, which is also 100 million.ca028c4
Dependencies¶
- Updated to Polars 2.0.
0a45c9c
2026-09-30¶
A quieter terminal¶
- Comparing every table in two PUDL builds prints hundreds of lines, nearly all of them for tables that did not change, and those buried the few that did.
The live table now lists only the tables that aren't identical, and leaves out the size columns, unless
--verboseis given. The progress counters, final summary and JSON report are unchanged, and--from-reportfollows the same flag. If every table is identical, a message says so instead of an empty table.e0ee36f - The status and primary key columns became emoji (✅, ⚠️, ❌ and 🔑, 🚫), which makes the table narrower so more of it fits on a screen.
2f10598
2026-09-26¶
Planning a notebook¶
- Wrote the plan for a Marimo notebook that visualizes a report: what it is for, how it should read, which technologies to consider, and what to show at the level of the dataset, the table, the column and the row.
The analysis is meant to live in library modules that are tested, and the notebook only presents it.
83bf513
2026-09-21¶
Stricter checks¶
- More of ruff's rules were turned on: unused arguments, implicit string concatenation, Pylint and tryceratops, and documenting every parameter.
Rules that don't suit the code, like limits on the number of arguments, were left off, with the reasons written down.
3ba80ce - Explicit
Anyis now an error in the type checker, since it is as good as no type at all but is counted as typed. The 16 uses were replaced, most notably by types for the parts of a datapackage descriptor that the tool reads, in a newdatapackagemodule.e4c0cca - The tests' docstrings are checked too, so every test says what it checks, and the more important or opaque ones say why.
7119780
GeoArrow compatibility¶
- Importing
pudl_diffin an environment where PUDL had already registered the GeoArrow WKB extension type (or the other way around) raised an error, because Polars refuses to register a type twice and can't say whether it is registered. A duplicate registration is now fine, and any other failure is still raised.50cc6c7
2026-09-20¶
Planning reports for nightly and stable builds¶
- Wrote down how PUDL's builds should use the tool: a nightly build compares itself against both the previous nightly build and the last stable release, a branch build against the last nightly, and a stable release against the previous release.
Reports are named for the two builds they compare, kept with the other outputs of the build, and published beside the nightly and stable data.
52feb84717b07a - A dataset can be read from one place and recorded in the report as another (
PudlDiffDataset(display_root=...)), so that a build's local outputs are described by the permanent location they will be deployed to.f2c14cc --from-reportshows a saved report again, as the same lines and summary a new comparison prints, without comparing anything.a8cde80
Becoming independent of PUDL¶
- The tool was to move out of the PUDL repository into a package of its own, so that it is useful outside of PUDL, doesn't add thousands of lines to PUDL's review, and can be installed by the scheduled jobs that archive every report. The first step was to make it stand alone while still in PUDL.
- Comparing row counts by partition, using the expressions in PUDL's dbt tests, was dropped.
It tied the tool to PUDL's repository layout, was brittle, and the per-partition breakdown was noisy.
The left-only and right-only Parquet files already show what is behind a change in the row count.
69b32dc - Logging uses the standard library, under loggers named
pudl_diff.<module>.b197db5 - A single module,
defaults, is the only one that knows anything about PUDL, and it does so lazily and optionally. It supplies the nightly build's location and the local outputs directory, and where a table's primary key comes from when its own datapackage doesn't say: PUDL's metadata if it's installed, else the last nightly build's datapackage, else none, with a warning. A test keeps every other module free of PUDL imports.28607a7 - The test fixtures moved out of PUDL's conftest, and the local
data/andreports/directories are ignored.f4193f35d72cae - The history was extracted with
git filter-repo, so it comes along. The imports were then renamed for the new package,pudl_diff, and its command line,pudl_diff.cli, and the tests moved totests/unit.b5ff178d39e600
A project of its own¶
- Packaging, hooks, CI, the devcontainer and the instructions for coding agents were adapted from Catalyst's Python template: hatchling with hatch-vcs, pixi environments and tasks, GitHub workflows for tests, docs, releases and lockfile updates.
b406d066b34c03 - The documentation was converted to Markdown for Zensical, and the reference for the report's fields became a page generated beside the JSON Schema, instead of a Sphinx page built from a template.
Zensical drops its default Markdown extensions when any are configured, so the configuration lists them all.
77c2595e37108f - Docstrings were converted from reStructuredText to Markdown, since both the API reference and the schema's descriptions are built from them, and neither is Sphinx.
A test keeps the Sphinx syntax out.
The schema also says
nullrather thannone.4e14e2c - The
LICENSEand the package metadata declare the license and that the package is typed, and the project has homepage and funding URLs.f9c299e223c869 - Allowed committing to
mainuntil the repository is on GitHub.946ed0c
Quality checks¶
- Type checking uses pyrefly at its strict preset instead of ty, following PUDL and the FERC XBRL extractor, and every part of
src/must be annotated.f4e9d7b - Test coverage must be 100%, counting branches, and is measured over the tests and scripts as well as the package, which finds tests that never run and helpers that nothing uses.
f4e9d7bcd4ce9b - Warnings are errors in the tests, and pytest's strict mode is on. Test output is a character per test, and every module's logger is mocked so that the expected failures don't print errors.
266b5494fe6ddd - The tests run on Linux, macOS and Windows, and the linters run once in CI.
f9c299e - Ruff adds to its default rules rather than replacing them, with PUDL's selection.
23979f0734ea61 - A development helper that summarizes pyrefly's coverage report lives in
scripts/, outside the published package.70d505f560b338 - Console script tests run the installed
pudl_diffon two small datasets, and more of the branches of the CLI, runner and summary are tested.38bf05e7c91f1c
Reports and behavior¶
- Reports record which version of
pudl_diffmade them, andpudl_diff --versionshows it. Reports will be archived and loaded long after they are made, so the version is required: a report that doesn't say is rejected, rather than credited to whatever happens to be installed. The schema version, which is that of the report's format, is unchanged, since no report has been published yet.c45bbf3 - PUDL stores geometries as GeoArrow WKB, an extension type Polars doesn't know, and every scan of a table with one warned.
It is now registered as the binary type it is stored as, so geometries are compared as the bytes they are. Other unknown types still warn.
abb8f04 - The temporary directory that holds the rows that differ (a
SpillDir) is deleted as soon as the report's Parquet files have been written, rather than when the result is garbage collected, which also raised a resource warning.7cec5ad
2026-09-19¶
One report for the whole comparison¶
- The tool wrote a JSON report for each table.
It now writes one
pudl_diff_report.jsonfor the whole run, built around aPudlDiffReportthat holds what pertains to the comparison as a whole: when it was made, the two datasets and their provenance, the options used, the tables in only one of them, and asummary. The tables are a dictionary ofTableDiffReports keyed by name, so a consumer can look one up directly, and the summary does the arithmetic across tables so consumers (and the terminal summary) don't have to. A run that fails before comparing anything, such as one where the datasets have no tables in common, still writes a report that says why.ce22e04511d561b1debf387c5ba8 - Pydantic models were chosen over dataclasses so that a report can be validated when it is loaded again for analysis or visualization.
- Each table's Parquet file size is recorded on both sides, with the difference and its percentage, because compression settings can change a file's size when its contents are the same.
511d561
A report that can be trusted and understood¶
is_identicalandsuccessused to be stored fields filled in by the code that built the report, so a report could contradict itself. They are now computed from the fields they summarize, and loading a report recomputes them rather than trusting the file.c47e016- Every field of the report is described, and the descriptions come from the docstrings, through the
ReportModelbase class of the models, so they are written once. The sections for rows with a primary key and without one are each either a summary or a record of why they were skipped, and astatusfield says which, so the schema can tell them apart.fafd686f3bf4a8 - The report's JSON Schema, generated by
report_json_schema(), is committed and published, and a generated reference page describes every field. A pre-commit hook regenerates them when the models change, and a test fails if they are out of date.683fa80 - Paths in the report no longer depend on where the tool was run: a local dataset is identified by its absolute path, and the Parquet files are named relative to the report, so the directory can be moved.
d5bd970
Splitting up a 2,400-line module¶
- The comparison code had grown into one module of about 2,400 lines, with a lot of logic in the CLI as well.
A notebook and perhaps a web app were to use it too, so it became a subpackage, with the CLI as a thin wrapper that parses options and dispatches.
The plan, including the dependency order of the modules, was checked for cycles before starting.
f84cd67 - To make the changes reviewable, code was moved verbatim first and edited in separate commits: first the pieces of the library, one module at a time (
formatting,dataset,schema,row_counts,rows,performance,table,outputsanddataset_report), then the presentation and orchestration code that lived in the CLI (terminalandrunner).19ee9e1337c6310940044fb39d342f7613170fdf8e3421806322c2bc2fee3e709072a719de37847fdc4d710fbb69a25eace97734ee909f21 - The edits then followed: names shared between modules lost their underscores, and
run_dataset_diffreplaced the loop that printed a line per table. It returns the report and announces progress through callbacks (whichTerminalProgresssupplies for the CLI), so that anything that isn't a terminal (a notebook, a web app) can run a comparison and get what the CLI writes.735f8065237804 - The layers of the package are documented, and the unit tests were split to mirror the modules, sharing the dataset fixtures with the CLI tests.
9550e45ea12a56896053b0ce5e2cf9d5b7b - Comparisons are meant to ignore the order of a table's rows and columns, but only a couple of cases had checked that.
Tests now shuffle both, with and without a primary key, including a composite one, and check that real changes are still found.
f7fa5c2 - The implementation notes were kept up to date along the way.
61210b4fb489264ff2ce6a883b5c1d0e958787a0e6
Terminal output¶
- Size changes are green when a table grew and red when it shrank, like the row and column counts, in place of blue and orange.
The column headings take two lines so the columns can be narrower, and the summary's counts of tables stand out.
4ad4d8e
2026-09-18¶
Defaults and documentation¶
- The nightly build is now the default
--left, the reference, and the local outputs are--right, so a diff reads as what has changed locally since the last nightly build, which is how the tool is mostly used.d503368 - Documented the tool, and added usage examples to
--help.1acc062 - Added short flags for the main options.
f87e57a
Bounding memory¶
- Row-level comparison materialized the joined tables, which doesn't scale.
The first change counted rows and changes with Polars' streaming engine, kept differing rows in temporary Parquet files instead of memory, and ran the join once instead of three times.
The option of returning pandas dataframes went away, since the results are lazy; callers can collect and convert.
2389de2 - Rows are now matched by a 64-bit hash.
Each table is reduced to one narrow row per hash of its primary key (or, with none, of the whole row with its floats quantized), and the two reductions are joined in a single streaming pass.
Full rows are read back only for the keys that differ.
For tables with a primary key, a hash of the exact non-key values finds the candidates for a change, which are then compared with the tolerance-aware logic, so a real change can't hide behind a quantization boundary.
9071765 - Tables without a primary key are compared as multisets, so a row that appears a different number of times is a difference, and duplicate primary keys are counted.
- A hash can collide, but the chance of masking a real change is about one in 10^11 per changed row at 300 million rows. On synthetic 30-million-row tables, peak memory fell from 12.4 GB to 4.1 GB with a primary key, and from 7.5 GB to 4.0 GB without.
Vocabulary¶
- The left table is the reference and the right is expected to differ from it, so the report says "change" instead of "mismatch", and the row count difference is right minus left.
A row diff always has both its keyed and unkeyed sections, and the one that doesn't apply says why instead of being null.
b77cd1e
Comparing many tables at once¶
- Comparing a whole dataset meant running the tool once per table, and each run paid to import all of PUDL.
It now takes any number of tables, or none for every table in both datasets, compares each in one process, and carries on if one fails.
The exit code is the highest of any table's.
1486f7e - The results are an aligned, colorized table, with a line per table as it finishes: its status, whether it has a primary key, columns and rows added, changed and removed in git-diff colors, the left size, and the time.
The end of the run totals them.
Log messages below the error level are hidden by default so they don't interrupt it.
69ba97f - Schema changes get the same treatment as row changes, with a count of columns, and the summary names the tables whose schema changed, since that can be disruptive to users.
b5d281c7726481
2026-09-17¶
The command line and its report¶
- Planned the CLI and its report as a series of reviewable tasks.
d66d533 - Each comparison records its time, peak memory and peak CPU, so reports can show what a comparison costs.
Memory is sampled from a background thread by a
PerformanceSampler, not read fromgetrusage, whose figure covers the process's whole life and has different units on macOS and Linux.b1e93d5 - Added the
pudl_diffcommand, which compares a table in two datasets, writes a JSON report and the Parquet files of the differing rows, and exits0for identical,1for different, and2if the comparison itself failed.7e163fc - A table's primary key can come from PUDL's own metadata when the dataset's datapackage is missing or stale, as happens in local development.
4c1416a
2026-09-16¶
All changes on this day (after the first commit)
A plan and the first comparisons¶
- Wrote down the plan: a shared, tested, documented tool that compares two PUDL outputs, on the principle that two tables are the same when they have the same columns and types, the same number of rows, and the same rows in any order, with floats compared as
numpy.isclosedoes. Polars was chosen because it can query remote Parquet lazily, with DuckDB as the fallback, and so was Click for the command line. The tool is for confirming that refactors and dependency upgrades didn't change the data, and for seeing what changed between builds.76b00e1 PudlDiffDatasetwraps a local or remote root and its datapackage, finds each table's primary key and Parquet file, and scans it lazily.f41aec5- Smoke tests against real PUDL outputs turned up two bugs with infinite values: two equal infinities compared as different, and opposite infinities compared as close or as the same row.
Both now match
numpy.isclose.a272d27 - The smoke tests also showed that comparing a table of hundreds of millions of rows to itself took 93 GB, and one of a billion rows was killed for running out of memory.
As a stopgap, row-level comparison is skipped over 100 million rows (
MAX_COMPARE_ROWS), leaving the schema and row counts, and a table whose rows were never compared is no longer reported as identical. The timings and memory of the runs were recorded to guide the real fixes, which came with the hashing of 2026-09-18 and the partitioning of 2026-10-06.