pudl_diff.table
¶
Comparing a single table between two datasets.
MAX_COMPARE_ROWS = 100000000
module-attribute
¶
The default of the command line's --max-compare-rows, which keeps a default
comparison from spending minutes on the very largest tables.
Row-level comparisons use Polars' streaming engine, spill their results to disk, and
join narrow 64-bit row hashes rather than whole rows, and tables of more than
MAX_ROWS_PER_PARTITION rows are compared in several passes, so memory doesn't limit
the size of table that can be compared; time does.
The API has no limit unless one is passed to compare_table(), which then skips
row-level comparison entirely for a larger table, while the cheaper schema and
row-count comparisons still run.
RowComparisonSkipReason = Literal['too_many_rows', 'incompatible_dtypes', 'mismatched_columns']
module-attribute
¶
Fixed set of reasons compare_table() skips the row-level comparison altogether
(see TableDiffResult.row_diff_skipped_reason).
A detailed, interpolated explanation (row counts, table names, etc.) is still logged
via logger.warning at the point of the skip; this only carries the category, so
report consumers can branch on it without parsing text.
TableDiffResult
dataclass
¶
The full comparison of a single table between two PUDL datasets.
Source code in src/pudl_diff/table.py
elapsed_seconds = 0.0
class-attribute
instance-attribute
¶
Wall-clock time compare_table() took to run this comparison.
is_identical
property
¶
Whether the table is functionally identical between the two datasets.
Conservatively False whenever row_diff is None, even if
the schema and row counts match: without a row-level comparison having
actually run, row content is unverified and can't be called identical.
peak_cpu_percent = 0.0
class-attribute
instance-attribute
¶
Peak per-interval CPU utilization observed while compare_table()
was running, as a percent of one core (e.g. 400.0 for four cores kept
fully busy at once). Sampled on the same background thread as
peak_rss_bytes; a rough gauge of how parallelized Polars' work
was, not an exact thread count.
peak_rss_bytes = 0
class-attribute
instance-attribute
¶
Peak resident set size attributable to this comparison: the highest
whole-process RSS observed while compare_table() was running, net
of the process's RSS just before it started. Sampled on a background
thread (see PerformanceSampler), so very short, sharp spikes
between samples may be missed.
right_table_name
instance-attribute
¶
Same as table_name unless a different right_table_name was
passed to compare_table(), e.g. to compare a core_ table against
the out_ table built from it.
row_diff
instance-attribute
¶
For a table with a primary key, computed over only the columns shared
by both datasets if their column sets differ (with a warning logged), since
the primary key still uniquely identifies rows either way. None if the
table has no primary key and the column sets differ - a row-level
comparison across changed columns isn't meaningful without one - if
either table has more than the configured row-level comparison cap, or if
the comparison didn't complete for some other reason; see
compare_table() and row_diff_skipped_reason.
row_diff_skipped_reason = None
class-attribute
instance-attribute
¶
Why row_diff is None: too many rows, incompatible column
dtypes, or changed columns without a usable primary key. None if
row-level comparison actually ran (whether or not it found any diffs).
TableDiffRun
dataclass
¶
The outcome of a run_table_diff() call.
Distinct from TableDiffResult.row_diff_skipped_reason, which
covers expected situations where row-level comparison is skipped but
the comparison otherwise completes normally (e.g. a table too large to
compare row-by-row). TableDiffRun instead covers the comparison
failing to complete at all - a missing datapackage, an S3 access
failure, or any other unexpected error - so a report can always be
produced even when compare_table() itself raises.
Source code in src/pudl_diff/table.py
compare_table(left, right, table_name, *, right_table_name=None, rtol=1e-05, atol=1e-08, max_compare_rows=None)
¶
Compare a single table between two PUDL datasets.
Times the comparison and samples this process's peak RSS and CPU
utilization while it runs (see PerformanceSampler), recording
all three on the returned TableDiffResult. See
_compare_table() for the comparison logic itself.
Source code in src/pudl_diff/table.py
row_diff_left_right_frames(row_diff)
¶
All differing rows, split by source table, in that table's own schema.
Returns a (lazy_frame, row_count) pair for each side.
For a table with a primary key, this is the symmetric difference of primary keys (rows present on only one side) plus, for shared keys, the left/right values of rows whose non-PK data differs. For a table without one, it's just the symmetric difference of whole rows.
Source code in src/pudl_diff/table.py
run_table_diff(left, right, table_name, *, right_table_name=None, rtol=1e-05, atol=1e-08, max_compare_rows=None)
¶
Run compare_table(), tolerating any failure it raises.
Takes the same arguments as compare_table() and passes them
through unchanged. Use this instead of calling compare_table()
directly when building a report that should always be produced, even if
the comparison itself blows up - e.g. because a dataset's
datapackage.json is missing, or an S3 root is unreachable.