pudl_diff.table_report
¶
The report on the comparison of one table.
Serializable summaries of the schema, row-count and row-level differences between two
tables, assembled into a TableDiffReport, and report_table_diff(), which
compares a table and builds its report.
RowDiffSectionSkipReason = Literal['too_many_rows', 'incompatible_dtypes', 'mismatched_columns', 'primary_key_available', 'no_primary_key']
module-attribute
¶
Reasons one section of the JSON report's row_diff (pk_diff or non_pk_diff)
has no summary.
The first three are the RowComparisonSkipReason values, for when the whole row-level
comparison was skipped; the last two mean that section simply doesn't apply, because
the table does or doesn't have a primary key.
DiffOptions
¶
Bases: ReportModel
The settings a dataset comparison was run with.
Recorded in the report because they affect how its results should be interpreted, e.g. whether a table's row-level comparison was skipped.
Source code in src/pudl_diff/table_report.py
atol = 1e-08
class-attribute
instance-attribute
¶
Absolute tolerance for float equality, as in numpy.isclose().
max_compare_rows = None
class-attribute
instance-attribute
¶
Row-level comparison is skipped for any table with more rows than this, or
None to compare tables however many rows they have.
max_output_rows = None
class-attribute
instance-attribute
¶
Cap on the rows written to each Parquet side-output file, or None to
write every differing row.
rtol = 1e-05
class-attribute
instance-attribute
¶
Relative tolerance for float equality, as in numpy.isclose().
NonPkRowDiffSummary
¶
Bases: ReportModel
The JSON-report form of a RowSetDiff without a primary key.
Source code in src/pudl_diff/table_report.py
is_identical
property
¶
Whether every row in one table has a matching row in the other.
multiplicity_changed_row_count
instance-attribute
¶
Number of distinct rows present in both tables, but a different number of times.
only_in_left_count
instance-attribute
¶
Number of rows only in the left table: removed rows. Rows are counted as a multiset, so if a row appears more times on the left, the surplus copies count.
only_in_right_count
instance-attribute
¶
Number of rows only in the right table: added rows, counted the same way.
status = 'compared'
class-attribute
instance-attribute
¶
Always "compared": this is what tells this apart from a skipped section.
symmetric_difference_count
instance-attribute
¶
only_in_left_count + only_in_right_count. Counts rows as a
multiset, so surplus copies of duplicated rows are included.
ParquetOutputSummary
¶
Bases: ReportModel
The JSON-report form of a single ParquetOutput.
Source code in src/pudl_diff/table_report.py
bytes
instance-attribute
¶
Size of the file in bytes.
hash
instance-attribute
¶
"sha256:<hexdigest>" of the file's contents, the convention PUDL's
datapackage.json uses for its own resource files.
path
instance-attribute
¶
Where the file is, relative to the directory that contains the report
(pudl_diff_report.json), so that the directory can be moved. Always written
with / separators.
size
property
¶
bytes in human-readable form, e.g. 12.3 MB.
from_parquet_output(output, report_dir)
classmethod
¶
Build from a ParquetOutput.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
output
|
ParquetOutput
|
The file that was written. |
required |
report_dir
|
str | PathLike[str]
|
The directory that the report will be in, which the file's path in the summary is relative to. |
required |
Source code in src/pudl_diff/table_report.py
PkRowDiffSummary
¶
Bases: ReportModel
The JSON-report form of a KeyedRowDiff.
Source code in src/pudl_diff/table_report.py
changed_row_count
instance-attribute
¶
Number of shared-primary-key rows with at least one differing non-primary-key value.
column_changes
instance-attribute
¶
Maps each non-primary-key column to the number of shared-primary-key rows where its value differs between the tables. Columns with no changes are omitted.
is_identical
property
¶
Whether the primary keys match and no shared-key row has changed.
only_in_left_count
instance-attribute
¶
Number of rows whose primary key is only in the left table: removed rows.
only_in_right_count
instance-attribute
¶
Number of rows whose primary key is only in the right table: added rows.
primary_key_columns
instance-attribute
¶
The names of the table's primary key columns.
primary_keys_identical
property
¶
Whether both tables have the same set of primary keys.
status = 'compared'
class-attribute
instance-attribute
¶
Always "compared": this is what tells this apart from a skipped section.
RowChanges
dataclass
¶
The row-level results of a table comparison, boiled down to counts.
Source code in src/pudl_diff/table_report.py
added = None
class-attribute
instance-attribute
¶
Rows only in the right table (for a table with a primary key, rows whose
primary key is only in the right table). None unless row-level comparison
ran.
changed = None
class-attribute
instance-attribute
¶
Rows whose primary key is in both tables but whose other values changed.
None unless row-level comparison ran on a table with a primary key.
has_primary_key = None
class-attribute
instance-attribute
¶
None if that isn't known, e.g. because the comparison failed.
removed = None
class-attribute
instance-attribute
¶
Like added, but for the left table.
skipped_reason = None
class-attribute
instance-attribute
¶
Why row-level comparison was skipped, if it was.
from_summary(row_diff)
classmethod
¶
Boil a table's row diff summary down to its counts.
Source code in src/pudl_diff/table_report.py
RowCountDiffSummary
¶
Bases: ReportModel
The JSON-report form of RowCountDiff.
Source code in src/pudl_diff/table_report.py
is_identical
property
¶
Whether the row counts match.
left_row_count
instance-attribute
¶
Total number of rows in the left table.
right_row_count
instance-attribute
¶
Total number of rows in the right table.
row_count_difference
instance-attribute
¶
right_row_count - left_row_count: the change in row count from the
reference (left) table to the right table.
from_row_count_diff(row_count_diff)
classmethod
¶
Build from a RowCountDiff.
Source code in src/pudl_diff/table_report.py
RowDiffSectionSkipped
¶
Bases: ReportModel
Stands in for a row diff summary that wasn't produced, and says why.
Source code in src/pudl_diff/table_report.py
skipped_reason
instance-attribute
¶
Why there is no summary in this section:
too_many_rows: either table has more rows thanDiffOptions.max_compare_rows, so no row-level comparison was made.incompatible_dtypes: the row-level comparison failed, most likely because the tables' columns have incompatible dtypes.mismatched_columns: the tables have different columns and no primary key, so their rows can't be compared meaningfully.primary_key_available: not skipped for a problem. The table has a primary key, so its comparison is inpk_diff, notnon_pk_diff.no_primary_key: likewise, the table has no primary key, so its comparison is innon_pk_diff, notpk_diff.
status = 'skipped'
class-attribute
instance-attribute
¶
Always "skipped": this is what tells this apart from a full summary.
RowDiffSummary
¶
Bases: ReportModel
The JSON-report form of row_diff.
Exactly one of pk_diff and non_pk_diff is a full summary
(unless row-level comparison was skipped entirely); the other is a
RowDiffSectionSkipped saying why it wasn't produced: no_primary_key
or primary_key_available when the table's primary key determined which
kind of comparison applies, or the reason the comparison was skipped
altogether (see RowComparisonSkipReason), in which case the one that would
have run carries that reason.
Source code in src/pudl_diff/table_report.py
left_only_parquet = None
class-attribute
instance-attribute
¶
The Parquet file of the rows found only in the left table (for a table with a
primary key, this includes the left-hand values of rows that changed), or None
if no file was written.
non_pk_diff
instance-attribute
¶
The row-level comparison of a table without a primary key: a full summary
(status is "compared"), or the reason there isn't one ("skipped").
pk_diff
instance-attribute
¶
The row-level comparison of a table with a primary key: a full summary
(status is "compared"), or the reason there isn't one ("skipped").
right_only_parquet = None
class-attribute
instance-attribute
¶
The same, for the right table.
SchemaDiffSummary
¶
Bases: ReportModel
The JSON-report form of SchemaDiff.
Source code in src/pudl_diff/table_report.py
columns_only_in_left
instance-attribute
¶
Names of the columns that are only in the left table: removed columns.
columns_only_in_right
instance-attribute
¶
Names of the columns that are only in the right table: added columns.
dtype_changes
instance-attribute
¶
Maps column name to a (left_dtype, right_dtype) pair of dtype
names, e.g. ("Int64", "Int32").
is_identical
property
¶
Whether the two schemas have the same columns and dtypes.
left_column_count
instance-attribute
¶
Number of columns in the left table.
right_column_count
instance-attribute
¶
Number of columns in the right table.
from_schema_diff(schema_diff)
classmethod
¶
Build from a SchemaDiff, stringifying its Polars dtypes.
Source code in src/pudl_diff/table_report.py
SizeComparison
¶
Bases: ReportModel
The sizes of the left and right side of a comparison, and how they differ.
Sizes are bytes on disk (or in cloud storage) of the Parquet file(s) being
compared. Everything but the two *_table_bytes fields is derived from them
when the report is serialized, and is None whenever either size is unknown.
Source code in src/pudl_diff/table_report.py
bytes_difference
property
¶
The change in size from the left to the right table.
right_table_bytes - left_table_bytes, so negative if the right side is
smaller. Compression changes show up here even if the contents don't.
bytes_difference_percent
property
¶
bytes_difference as a percentage of left_table_bytes.
None if the left size is unknown or zero.
bytes_difference_size
property
¶
bytes_difference in human-readable form, e.g. -1.2 MB.
left_table_bytes = None
class-attribute
instance-attribute
¶
Size in bytes of the left table's Parquet file(s), or None if unknown.
left_table_size
property
¶
left_table_bytes in human-readable form, e.g. 12.3 MB.
right_table_bytes = None
class-attribute
instance-attribute
¶
Size in bytes of the right table's Parquet file(s), or None if unknown.
right_table_size
property
¶
right_table_bytes in human-readable form.
TableDiffReport
¶
Bases: SizeComparison
A single table comparison, in the form saved in the PUDL Diff JSON report.
Built by build_table_diff_report() from a TableDiffRun, and
one entry in tables. Fields that describe the whole
comparison of the two datasets (when it was run, the datasets' provenance) live
on the PudlDiffReport instead. Contains no row-level data itself -
only counts and summaries; the actual differing rows are written
separately as Parquet files (see write_row_diff_parquet()) and
referenced from row_diff.
Source code in src/pudl_diff/table_report.py
442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 | |
elapsed_seconds = None
class-attribute
instance-attribute
¶
Wall-clock time the comparison of this table took, or None if it failed.
error = None
class-attribute
instance-attribute
¶
Exception message plus traceback, if the comparison failed to
complete. None if success is True.
is_identical
property
¶
Whether the table is functionally identical between the two datasets.
Conservatively False whenever success is False, since a
failed comparison can't establish that the tables are identical, and
whenever the row-level comparison didn't run (it was skipped), since then
the rows are unverified even if the schema and row counts match.
left_table_name
instance-attribute
¶
The name of the table in the left dataset.
left_table_path
instance-attribute
¶
The path or URL of the table's Parquet file in the left dataset. Worked out from the dataset's root and the table's name, so it is given even if the file doesn't exist, e.g. because the comparison failed. Absolute, for a dataset on the local filesystem.
peak_cpu_percent = None
class-attribute
instance-attribute
¶
The highest CPU utilization sampled during this comparison, as a percentage of
one core: 400.0 means four cores kept fully busy. A rough gauge of how
parallel the work was. None if the comparison failed.
peak_rss
property
¶
peak_rss_bytes in human-readable form, e.g. 1.2 GB.
peak_rss_bytes = None
class-attribute
instance-attribute
¶
The most memory (resident set size) the process used during this comparison
beyond what it was using when the comparison started, in bytes. Sampled, so a
very short spike could be missed. None if the comparison failed.
right_table_name
instance-attribute
¶
The name of the table in the right dataset. Differs from
left_table_name only when two differently named tables were compared,
e.g. a core_ table against the out_ table built from it.
right_table_path
instance-attribute
¶
The path or URL of the table's Parquet file in the right dataset. Absolute, for a dataset on the local filesystem.
row_count_diff = None
class-attribute
instance-attribute
¶
How the tables' row counts differ, or None if the comparison failed.
row_diff = None
class-attribute
instance-attribute
¶
How the tables' rows differ, or None if the comparison failed. If the
row-level comparison was skipped, its sections say why.
schema_diff = None
class-attribute
instance-attribute
¶
How the tables' columns and dtypes differ, or None if the comparison
failed.
success
property
¶
Whether the comparison completed at all, successfully or not.
See TableDiffRun. Distinct from is_identical: a
comparison can succeed and still find the tables different.
build_table_diff_report(run, left, right, table_name, *, right_table_name=None, parquet_outputs=None, report_dir=None)
¶
Build the JSON-report form of a table comparison.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
run
|
TableDiffRun
|
The comparison's outcome, from |
required |
left
|
PudlDiffDataset
|
The "left" dataset that was compared. |
required |
right
|
PudlDiffDataset
|
The "right" dataset compared against it. |
required |
table_name
|
str
|
Name of the table compared in |
required |
right_table_name
|
str | None
|
Name of the table compared in |
None
|
parquet_outputs
|
RowDiffParquetOutputs | None
|
The Parquet side-output files written for this
table's row diff, from |
None
|
report_dir
|
str | PathLike[str] | None
|
The directory that the report will be written to, which the
paths of the |
None
|
Source code in src/pudl_diff/table_report.py
report_table_diff(left, right, table_name, output_path, *, right_table_name=None, options=None)
¶
Compare a table, write its Parquet side-outputs, and report on it.
Never raises because the comparison failed: see run_table_diff().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left
|
PudlDiffDataset
|
The "left" dataset to compare. |
required |
right
|
PudlDiffDataset
|
The "right" dataset to compare against |
required |
table_name
|
str
|
Name of the table to compare in |
required |
output_path
|
str | PathLike[str]
|
Directory to write the differing rows' Parquet files into. It's also where the report is to be written: the files' paths in the report are relative to it. |
required |
right_table_name
|
str | None
|
Name of the table to compare in |
None
|
options
|
DiffOptions | None
|
How to run the comparison. Defaults to |
None
|