pudl_diff.runner
¶
Comparing many tables between two datasets.
run_dataset_diff(left, right, output_path, *, table_names=(), right_table=None, options=None, on_tables_resolved=None, on_table_compared=None)
¶
Compare tables between two datasets, one at a time, and report on them all.
A table whose comparison fails doesn't stop the others being compared, and if
the tables to compare can't even be determined, the returned report records why
in its error instead of raising.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left
|
PudlDiffDataset
|
The "left" dataset. |
required |
right
|
PudlDiffDataset
|
The "right" dataset, compared against it. |
required |
output_path
|
Path
|
Directory to write each differing table's Parquet outputs into. |
required |
table_names
|
Sequence[str]
|
The tables to compare. If empty, every table with a Parquet file in both datasets is compared. |
()
|
right_table
|
str | None
|
The name of the table to compare against in |
None
|
options
|
DiffOptions | None
|
The settings to compare with. |
None
|
on_tables_resolved
|
Callable[[list[str]], None] | None
|
Called once with the names of the tables that are about to be compared, before any of them is. |
None
|
on_table_compared
|
Callable[[str, TableDiffReport], None] | None
|
Called with each table's name and report as soon as it's been compared. |
None
|