Skip to content

PUDL Diff Report Schema

pudl_diff writes a JSON report of its comparison, described in Usage. This page is the reference for every field of that report. It is generated from the report's Pydantic models, which are also where the descriptions below are written, so it always matches the code.

The report is described by a JSON Schema, which you can use to validate a report, or to give a program or a coding agent a machine-readable description of it: pudl_diff_report.schema.json. The same schema is available in Python, and the report can be loaded and checked with Pydantic:

from pudl_diff.dataset_report import PudlDiffReport
from pudl_diff.report_schema import report_json_schema

schema = report_json_schema()
report = PudlDiffReport.model_validate_json(report_path.read_text(encoding="utf-8"))

Every field listed here is always present in a report, even if its value is null. A field marked derived is calculated from the others: it is written to the report for convenience, but ignored when a report is loaded. The version of the report format is in its schema_version field, which follows the major.minor.patch convention, and the version of pudl_diff that made the report is in its pudl_diff_version.

The copies of the schema and of this page in the repository are kept up to date by the pudl-diff-schema pre-commit hook, which rewrites them when the code in pudl_diff changes, and checked by a unit test. To update them by hand, run pixi run python -m pudl_diff.report_schema.

PudlDiffReport

The full comparison of two PUDL datasets: the saved JSON report.

Built by build_pudl_diff_report(). Holds everything that pertains to the comparison as a whole, plus a TableDiffReport for each table.

schema_version

Type: string.

Version of this report format, in major.minor.patch form.

pudl_diff_version

Type: string.

Version of the pudl_diff package that made this report, e.g. 0.1.0, or for a development build, a version that says which commit it was built from. Unlike schema_version, which only changes when the report's format does, this changes with every release, so it records exactly which code produced the report. It has no default, so a report never claims to be from a version that didn't write it.

created

Type: string.

UTC ISO-8601 timestamp of when this report was generated.

elapsed_seconds

Type: number or null.

Wall-clock time the whole comparison took.

left_dataset

Type: DatasetInfo.

The reference dataset, e.g. the last nightly build. Additions, removals and changes are all measured from it to the right dataset.

right_dataset

Type: DatasetInfo.

The dataset that was compared against the left one, e.g. a local build.

options

Type: DiffOptions.

The settings the comparison was run with.

tables_only_in_left

Type: array of string.

Tables found only in the left dataset, which aren't compared.

tables_only_in_right

Type: array of string.

Tables found only in the right dataset, which aren't compared.

summary

Type: PudlDiffSummary.

Totals over all the compared tables.

tables

Type: object mapping names to TableDiffReport.

Each compared table's report, keyed by its name in the left dataset.

error

Type: string or null.

Why the comparison as a whole failed, e.g. no tables could be listed. This is null when the only failures are of individual tables, which each record their own error.

success

Type: boolean. Derived from the other fields.

Whether the comparison completed.

That is, error is null and so is every table's. Distinct from is_identical: a comparison can succeed and still find differences.

is_identical

Type: boolean. Derived from the other fields.

Whether every compared table is identical, and success is true.

Tables found in only one dataset don't count against this.

DatasetInfo

One of the two compared datasets: where it is, and where it came from.

id

Type: string or null.

The dataset's build UUID.

created

Type: string or null.

UTC ISO-8601 timestamp of when this dataset was built - distinct from the report's own created, which is when the comparison was run.

git_sha

Type: string or null.

The git commit SHA of the PUDL code that built the dataset.

git_tags

Type: array of string or null.

The git tags on that commit, e.g. release versions like v2026.1.0.

root

Type: string.

The root path or URL of the dataset's Parquet files. For a dataset on the local filesystem, an absolute path with any symlinks resolved, so it doesn't depend on the directory the comparison was run from.

DiffOptions

The settings a dataset comparison was run with.

Recorded in the report because they affect how its results should be interpreted, e.g. whether a table's row-level comparison was skipped.

rtol

Type: number.

Relative tolerance for float equality, as in numpy.isclose().

atol

Type: number.

Absolute tolerance for float equality, as in numpy.isclose().

max_compare_rows

Type: integer or null.

Row-level comparison is skipped for any table with more rows than this, or null to compare tables however many rows they have.

max_output_rows

Type: integer or null.

Cap on the rows written to each Parquet side-output file, or null to write every differing row.

PudlDiffSummary

Totals over every table in a PudlDiffReport.

Saves consumers from aggregating the tables themselves.

The size fields (see SizeComparison) total only the tables whose size is known on both sides.

left_table_bytes

Type: integer or null.

Size in bytes of the left table's Parquet file(s), or null if unknown.

right_table_bytes

Type: integer or null.

Size in bytes of the right table's Parquet file(s), or null if unknown.

table_count

Type: integer.

Number of tables compared, including any whose comparison failed.

identical_table_count

Type: integer.

Tables whose comparison completed and found no differences.

changed_table_count

Type: integer.

Tables whose comparison completed and found differences.

failed_table_count

Type: integer.

Tables whose comparison failed to complete.

failed_tables

Type: array of string.

The names of the tables whose comparison failed.

schema_changed_tables

Type: array of string.

Tables with columns added or removed, or with changed dtypes.

left_row_count

Type: integer.

Total rows in the left tables, over all tables that could be counted.

right_row_count

Type: integer.

Total rows in the right tables, over all tables that could be counted.

rows_added

Type: integer.

Rows only in the right table, summed over tables with a row-level comparison.

rows_changed

Type: integer.

Rows with the same primary key but changed values, summed over tables with a row-level comparison and a primary key.

rows_removed

Type: integer.

Rows only in the left table, summed over tables with a row-level comparison.

no_row_diff_table_count

Type: integer.

Tables with no row-level comparison, whether skipped or failed. Their rows count towards the row totals, but not the rows added, changed or removed.

no_row_diff_left_row_count

Type: integer.

Total rows in the left side of those tables.

columns_added

Type: integer.

Columns only in the right table, summed over all the tables.

columns_changed

Type: integer.

Shared columns whose dtype changed, summed over all the tables.

columns_removed

Type: integer.

Columns only in the left table, summed over all the tables.

peak_rss_bytes

Type: integer or null.

The highest peak_rss_bytes of any table.

peak_rss_table

Type: string or null.

The table with that peak memory use.

left_table_size

Type: string or null. Derived from the other fields.

left_table_bytes in human-readable form, e.g. 12.3 MB.

right_table_size

Type: string or null. Derived from the other fields.

right_table_bytes in human-readable form.

bytes_difference

Type: integer or null. Derived from the other fields.

The change in size from the left to the right table.

right_table_bytes - left_table_bytes, so negative if the right side is smaller. Compression changes show up here even if the contents don't.

bytes_difference_size

Type: string or null. Derived from the other fields.

bytes_difference in human-readable form, e.g. -1.2 MB.

bytes_difference_percent

Type: number or null. Derived from the other fields.

bytes_difference as a percentage of left_table_bytes.

null if the left size is unknown or zero.

peak_rss

Type: string or null. Derived from the other fields.

peak_rss_bytes in human-readable form, e.g. 1.2 GB.

TableDiffReport

A single table comparison, in the form saved in the PUDL Diff JSON report.

Built by build_table_diff_report() from a TableDiffRun, and one entry in tables. Fields that describe the whole comparison of the two datasets (when it was run, the datasets' provenance) live on the PudlDiffReport instead. Contains no row-level data itself - only counts and summaries; the actual differing rows are written separately as Parquet files (see write_row_diff_parquet()) and referenced from row_diff.

left_table_bytes

Type: integer or null.

Size in bytes of the left table's Parquet file(s), or null if unknown.

right_table_bytes

Type: integer or null.

Size in bytes of the right table's Parquet file(s), or null if unknown.

left_table_name

Type: string.

The name of the table in the left dataset.

left_table_path

Type: string.

The path or URL of the table's Parquet file in the left dataset. Worked out from the dataset's root and the table's name, so it is given even if the file doesn't exist, e.g. because the comparison failed. Absolute, for a dataset on the local filesystem.

right_table_name

Type: string.

The name of the table in the right dataset. Differs from left_table_name only when two differently named tables were compared, e.g. a core_ table against the out_ table built from it.

right_table_path

Type: string.

The path or URL of the table's Parquet file in the right dataset. Absolute, for a dataset on the local filesystem.

elapsed_seconds

Type: number or null.

Wall-clock time the comparison of this table took, or null if it failed.

peak_rss_bytes

Type: integer or null.

The most memory (resident set size) the process used during this comparison beyond what it was using when the comparison started, in bytes. Sampled, so a very short spike could be missed. null if the comparison failed.

peak_cpu_percent

Type: number or null.

The highest CPU utilization sampled during this comparison, as a percentage of one core: 400.0 means four cores kept fully busy. A rough gauge of how parallel the work was. null if the comparison failed.

schema_diff

Type: SchemaDiffSummary or null.

How the tables' columns and dtypes differ, or null if the comparison failed.

row_count_diff

Type: RowCountDiffSummary or null.

How the tables' row counts differ, or null if the comparison failed.

row_diff

Type: RowDiffSummary or null.

How the tables' rows differ, or null if the comparison failed. If the row-level comparison was skipped, its sections say why.

error

Type: string or null.

Exception message plus traceback, if the comparison failed to complete. null if success is true.

left_table_size

Type: string or null. Derived from the other fields.

left_table_bytes in human-readable form, e.g. 12.3 MB.

right_table_size

Type: string or null. Derived from the other fields.

right_table_bytes in human-readable form.

bytes_difference

Type: integer or null. Derived from the other fields.

The change in size from the left to the right table.

right_table_bytes - left_table_bytes, so negative if the right side is smaller. Compression changes show up here even if the contents don't.

bytes_difference_size

Type: string or null. Derived from the other fields.

bytes_difference in human-readable form, e.g. -1.2 MB.

bytes_difference_percent

Type: number or null. Derived from the other fields.

bytes_difference as a percentage of left_table_bytes.

null if the left size is unknown or zero.

success

Type: boolean. Derived from the other fields.

Whether the comparison completed at all, successfully or not.

See TableDiffRun. Distinct from is_identical: a comparison can succeed and still find the tables different.

is_identical

Type: boolean. Derived from the other fields.

Whether the table is functionally identical between the two datasets.

Conservatively false whenever success is false, since a failed comparison can't establish that the tables are identical, and whenever the row-level comparison didn't run (it was skipped), since then the rows are unverified even if the schema and row counts match.

peak_rss

Type: string or null. Derived from the other fields.

peak_rss_bytes in human-readable form, e.g. 1.2 GB.

SchemaDiffSummary

The JSON-report form of SchemaDiff.

columns_only_in_left

Type: array of string.

Names of the columns that are only in the left table: removed columns.

columns_only_in_right

Type: array of string.

Names of the columns that are only in the right table: added columns.

dtype_changes

Type: object mapping names to [string, string].

Maps column name to a (left_dtype, right_dtype) pair of dtype names, e.g. ("Int64", "Int32").

left_column_count

Type: integer.

Number of columns in the left table.

right_column_count

Type: integer.

Number of columns in the right table.

is_identical

Type: boolean. Derived from the other fields.

Whether the two schemas have the same columns and dtypes.

RowCountDiffSummary

The JSON-report form of RowCountDiff.

left_row_count

Type: integer.

Total number of rows in the left table.

right_row_count

Type: integer.

Total number of rows in the right table.

row_count_difference

Type: integer.

right_row_count - left_row_count: the change in row count from the reference (left) table to the right table.

is_identical

Type: boolean. Derived from the other fields.

Whether the row counts match.

RowDiffSummary

The JSON-report form of row_diff.

Exactly one of pk_diff and non_pk_diff is a full summary (unless row-level comparison was skipped entirely); the other is a RowDiffSectionSkipped saying why it wasn't produced: no_primary_key or primary_key_available when the table's primary key determined which kind of comparison applies, or the reason the comparison was skipped altogether (see RowComparisonSkipReason), in which case the one that would have run carries that reason.

pk_diff

Type: PkRowDiffSummary or RowDiffSectionSkipped, told apart by status.

The row-level comparison of a table with a primary key: a full summary (status is "compared"), or the reason there isn't one ("skipped").

non_pk_diff

Type: NonPkRowDiffSummary or RowDiffSectionSkipped, told apart by status.

The row-level comparison of a table without a primary key: a full summary (status is "compared"), or the reason there isn't one ("skipped").

left_only_parquet

Type: ParquetOutputSummary or null.

The Parquet file of the rows found only in the left table (for a table with a primary key, this includes the left-hand values of rows that changed), or null if no file was written.

right_only_parquet

Type: ParquetOutputSummary or null.

The same, for the right table.

PkRowDiffSummary

The JSON-report form of a KeyedRowDiff.

status

Type: "compared".

Always "compared": this is what tells this apart from a skipped section.

primary_key_columns

Type: array of string.

The names of the table's primary key columns.

only_in_left_count

Type: integer.

Number of rows whose primary key is only in the left table: removed rows.

only_in_right_count

Type: integer.

Number of rows whose primary key is only in the right table: added rows.

changed_row_count

Type: integer.

Number of shared-primary-key rows with at least one differing non-primary-key value.

column_changes

Type: object mapping names to integer.

Maps each non-primary-key column to the number of shared-primary-key rows where its value differs between the tables. Columns with no changes are omitted.

primary_keys_identical

Type: boolean. Derived from the other fields.

Whether both tables have the same set of primary keys.

is_identical

Type: boolean. Derived from the other fields.

Whether the primary keys match and no shared-key row has changed.

RowDiffSectionSkipped

Stands in for a row diff summary that wasn't produced, and says why.

status

Type: "skipped".

Always "skipped": this is what tells this apart from a full summary.

skipped_reason

Type: "too_many_rows" or "incompatible_dtypes" or "mismatched_columns" or "primary_key_available" or "no_primary_key".

Why there is no summary in this section:

  • too_many_rows: either table has more rows than DiffOptions.max_compare_rows, so no row-level comparison was made.
  • incompatible_dtypes: the row-level comparison failed, most likely because the tables' columns have incompatible dtypes.
  • mismatched_columns: the tables have different columns and no primary key, so their rows can't be compared meaningfully.
  • primary_key_available: not skipped for a problem. The table has a primary key, so its comparison is in pk_diff, not non_pk_diff.
  • no_primary_key: likewise, the table has no primary key, so its comparison is in non_pk_diff, not pk_diff.

NonPkRowDiffSummary

The JSON-report form of a RowSetDiff without a primary key.

status

Type: "compared".

Always "compared": this is what tells this apart from a skipped section.

only_in_left_count

Type: integer.

Number of rows only in the left table: removed rows. Rows are counted as a multiset, so if a row appears more times on the left, the surplus copies count.

only_in_right_count

Type: integer.

Number of rows only in the right table: added rows, counted the same way.

symmetric_difference_count

Type: integer.

only_in_left_count + only_in_right_count. Counts rows as a multiset, so surplus copies of duplicated rows are included.

multiplicity_changed_row_count

Type: integer.

Number of distinct rows present in both tables, but a different number of times.

is_identical

Type: boolean. Derived from the other fields.

Whether every row in one table has a matching row in the other.

ParquetOutputSummary

The JSON-report form of a single ParquetOutput.

path

Type: string.

Where the file is, relative to the directory that contains the report (pudl_diff_report.json), so that the directory can be moved. Always written with / separators.

bytes

Type: integer.

Size of the file in bytes.

hash

Type: string.

"sha256:<hexdigest>" of the file's contents, the convention PUDL's datapackage.json uses for its own resource files.

size

Type: string. Derived from the other fields.

bytes in human-readable form, e.g. 12.3 MB.