Is there an existing issue for this?
Current Behavior
When operational metadata is enabled (features.operationalMetadataEnabled, the default) the framework adds the operational metadata column(s) (e.g. meta_load_details with record_insert_timestamp / pipeline_update_id) to every row read. CDCFlow.create and CDCSnapshotFlow then pass the spec's track_history_except_column_list to dp.create_auto_cdc_flow / dp.create_auto_cdc_from_snapshot_flow unchanged, so for scd_type: "2" AUTO CDC tracks history on the operational metadata too.
The metadata changes on every load, so a row that is re-delivered with identical business values always opens a new SCD2 version, and the old one is closed. Excluding the file metadata columns via except_column_list does not help, because the operational metadata is added by the framework and still differs.
Expected Behavior
Operational metadata should not drive SCD2 history. A re-delivered row whose tracked business columns are unchanged should not open a new version. The metadata column should still be written and updated on the current row.
Steps To Reproduce
- An SCD2 dataflow over a file source, with
cdc_settings: {keys: [id], sequence_by: <file modification time>, scd_type: "2"} and the file metadata columns in except_column_list.
- Land a file containing one row and run the pipeline.
- Land an identical copy of that file under a new name and run the pipeline again.
- The row now has two versions. The only column that differs between them is
meta_load_details.
Channel
CURRENT
Relevant log output
Version comparison for the re-delivered key (business columns identical, only the metadata differs):
SAME <all business columns>
DIFF meta_load_details | record_insert_timestamp 03:43:44 ... | record_insert_timestamp 03:45:31 ...
Found while end-to-end testing our derivative of the framework (v0.24.1). A PR with a fix and unit tests follows.
Is there an existing issue for this?
Current Behavior
When operational metadata is enabled (
features.operationalMetadataEnabled, the default) the framework adds the operational metadata column(s) (e.g.meta_load_detailswithrecord_insert_timestamp/pipeline_update_id) to every row read.CDCFlow.createandCDCSnapshotFlowthen pass the spec'strack_history_except_column_listtodp.create_auto_cdc_flow/dp.create_auto_cdc_from_snapshot_flowunchanged, so forscd_type: "2"AUTO CDC tracks history on the operational metadata too.The metadata changes on every load, so a row that is re-delivered with identical business values always opens a new SCD2 version, and the old one is closed. Excluding the file metadata columns via
except_column_listdoes not help, because the operational metadata is added by the framework and still differs.Expected Behavior
Operational metadata should not drive SCD2 history. A re-delivered row whose tracked business columns are unchanged should not open a new version. The metadata column should still be written and updated on the current row.
Steps To Reproduce
cdc_settings: {keys: [id], sequence_by: <file modification time>, scd_type: "2"}and the file metadata columns inexcept_column_list.meta_load_details.Channel
CURRENT
Relevant log output
Found while end-to-end testing our derivative of the framework (v0.24.1). A PR with a fix and unit tests follows.