Stale Customer Profiles After an Incremental Load

PySpark · incremental processing · Intermediate · about 25 minutes

Impact

After the incremental change-data-capture (CDC) load, some customer profiles show an older value even though the change feed delivered a newer one.

Symptoms

Evidence

cdc_data rows for customer_id = 1

customer_idnameageupdate_time
1Alice312024-01-02 09:00:00
1Alice322024-01-02 11:00:00

Row counts, last run

datasetrows
main_data3
cdc_data (this batch)4
distinct customer_id in cdc_data3
published table3

Freshness check, customer 1

measurevalue
latest update_time in the batch2024-01-02 11:00:00
update_time of the published value2024-01-02 09:00:00

Current collapse step (excerpt)

# keep one change per customer before merging into main_data
latest = cdc_data.dropDuplicates(['customer_id'])

Investigation task

Apply the CDC batch to main_data so each customer_id ends up with exactly one row carrying its most recent change. Build df_result with customer_id, name and age, and finish with df_result.show().

The fix is written and graded in the regular challenge editor. The postmortem unlocks once the incident is resolved.

All production incidents