pudl_diff.dataset_report
¶
The dataset-level report, summarizing all the tables compared.
REPORT_FILENAME = 'pudl_diff_report.json'
module-attribute
¶
Name of the JSON report, written to the output directory.
REPORT_SCHEMA_VERSION = '1.1.0'
module-attribute
¶
Version of the JSON report format written by build_pudl_diff_report().
DatasetInfo
¶
Bases: DatasetProvenance
One of the two compared datasets: where it is, and where it came from.
Source code in src/pudl_diff/dataset_report.py
root
instance-attribute
¶
The root path or URL of the dataset's Parquet files. For a dataset on the local filesystem, an absolute path with any symlinks resolved, so it doesn't depend on the directory the comparison was run from.
PudlDiffReport
¶
Bases: ReportModel
The full comparison of two PUDL datasets: the saved JSON report.
Built by build_pudl_diff_report(). Holds everything that pertains to the
comparison as a whole, plus a TableDiffReport for each table.
Source code in src/pudl_diff/dataset_report.py
created
instance-attribute
¶
UTC ISO-8601 timestamp of when this report was generated.
elapsed_seconds = None
class-attribute
instance-attribute
¶
Wall-clock time the whole comparison took.
error = None
class-attribute
instance-attribute
¶
Why the comparison as a whole failed, e.g. no tables could be listed. This
is None when the only failures are of individual tables, which each
record their own error.
exit_code
property
¶
The CLI's exit status for this report.
0 if everything is identical, 1 if any differ, 2 if any
comparison failed.
is_identical
property
¶
Whether every compared table is identical, and success is True.
Tables found in only one dataset don't count against this.
left_dataset
instance-attribute
¶
The reference dataset, e.g. the last nightly build. Additions, removals and changes are all measured from it to the right dataset.
options
instance-attribute
¶
The settings the comparison was run with.
pudl_diff_version
instance-attribute
¶
Version of the pudl_diff package that made this report, e.g. 0.1.0, or for
a development build, a version that says which commit it was built from. Unlike
schema_version, which only changes when the report's format does, this changes
with every release, so it records exactly which code produced the report. It has
no default, so a report never claims to be from a version that didn't write it.
right_dataset
instance-attribute
¶
The dataset that was compared against the left one, e.g. a local build.
schema_version = REPORT_SCHEMA_VERSION
class-attribute
instance-attribute
¶
Version of this report format, in major.minor.patch form.
success
property
¶
Whether the comparison completed.
That is, error is None and so is every table's. Distinct from
is_identical: a comparison can succeed and still find differences.
summary
instance-attribute
¶
Totals over all the compared tables.
tables
instance-attribute
¶
Each compared table's report, keyed by its name in the left dataset.
tables_only_in_left
instance-attribute
¶
Tables found only in the left dataset, which aren't compared.
tables_only_in_right
instance-attribute
¶
Tables found only in the right dataset, which aren't compared.
PudlDiffSummary
¶
Bases: SizeComparison
Totals over every table in a PudlDiffReport.
Saves consumers from aggregating the tables themselves.
The size fields (see SizeComparison) total only the tables whose size
is known on both sides.
Source code in src/pudl_diff/dataset_report.py
38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 | |
changed_table_count
instance-attribute
¶
Tables whose comparison completed and found differences.
columns_added
instance-attribute
¶
Columns only in the right table, summed over all the tables.
columns_changed
instance-attribute
¶
Shared columns whose dtype changed, summed over all the tables.
columns_removed
instance-attribute
¶
Columns only in the left table, summed over all the tables.
failed_table_count
instance-attribute
¶
Tables whose comparison failed to complete.
failed_tables
instance-attribute
¶
The names of the tables whose comparison failed.
identical_table_count
instance-attribute
¶
Tables whose comparison completed and found no differences.
left_row_count
instance-attribute
¶
Total rows in the left tables, over all tables that could be counted.
no_row_diff_left_row_count
instance-attribute
¶
Total rows in the left side of those tables.
no_row_diff_table_count
instance-attribute
¶
Tables with no row-level comparison, whether skipped or failed. Their rows count towards the row totals, but not the rows added, changed or removed.
peak_rss
property
¶
peak_rss_bytes in human-readable form, e.g. 1.2 GB.
peak_rss_bytes = None
class-attribute
instance-attribute
¶
The highest peak_rss_bytes of any table.
peak_rss_table = None
class-attribute
instance-attribute
¶
The table with that peak memory use.
right_row_count
instance-attribute
¶
Total rows in the right tables, over all tables that could be counted.
rows_added
instance-attribute
¶
Rows only in the right table, summed over tables with a row-level comparison.
rows_changed
instance-attribute
¶
Rows with the same primary key but changed values, summed over tables with a row-level comparison and a primary key.
rows_removed
instance-attribute
¶
Rows only in the left table, summed over tables with a row-level comparison.
schema_changed_tables
instance-attribute
¶
Tables with columns added or removed, or with changed dtypes.
table_count
instance-attribute
¶
Number of tables compared, including any whose comparison failed.
from_tables(tables)
classmethod
¶
Total up the reports of individual tables, keyed by table name.
Source code in src/pudl_diff/dataset_report.py
TableOutcome
dataclass
¶
What happened when comparing one table, for display.
Source code in src/pudl_diff/dataset_report.py
columns_added = None
class-attribute
instance-attribute
¶
Columns only in the right table. None if the comparison failed.
columns_removed = None
class-attribute
instance-attribute
¶
Columns only in the left table.
dtypes_changed = 0
class-attribute
instance-attribute
¶
Number of shared columns whose dtype differs between the tables.
exit_code
instance-attribute
¶
0 if identical, 1 if different, 2 if the comparison failed.
build_pudl_diff_report(left, right, tables, *, options=None, tables_only_in_left=(), tables_only_in_right=(), elapsed_seconds=None, error=None)
¶
Assemble the report on a comparison of two whole datasets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left
|
PudlDiffDataset
|
The "left" dataset that was compared. |
required |
right
|
PudlDiffDataset
|
The "right" dataset compared against it. |
required |
tables
|
dict[str, TableDiffReport]
|
The report on each compared table, keyed by its name in |
required |
options
|
DiffOptions | None
|
The settings the comparison ran with. |
None
|
tables_only_in_left
|
Iterable[str]
|
Tables that weren't compared as they aren't in
|
()
|
tables_only_in_right
|
Iterable[str]
|
Likewise, for tables not in |
()
|
elapsed_seconds
|
float | None
|
How long the whole comparison took. |
None
|
error
|
str | None
|
Why the comparison failed as a whole, if it did - e.g. because no tables could be found to compare. |
None
|
Source code in src/pudl_diff/dataset_report.py
table_outcome(table_name, report)
¶
Boil a table's report down to what we display.