Stale Customer Profiles After an Incremental Load
PySpark · incremental processing · Intermediate · about 25 minutes
Impact
After the incremental change-data-capture (CDC) load, some customer profiles show an older value even though the change feed delivered a newer one.
Symptoms
- The incremental job finished with status SUCCESS.
- Customer 1 shows age 31, but the latest change for that customer says 32.
- Only customers with more than one change in the same batch are affected.
- A full reload from source shows the correct value.
Evidence
cdc_data rows for customer_id = 1
| customer_id | name | age | update_time |
|---|---|---|---|
| 1 | Alice | 31 | 2024-01-02 09:00:00 |
| 1 | Alice | 32 | 2024-01-02 11:00:00 |
Row counts, last run
| dataset | rows |
|---|---|
| main_data | 3 |
| cdc_data (this batch) | 4 |
| distinct customer_id in cdc_data | 3 |
| published table | 3 |
Freshness check, customer 1
| measure | value |
|---|---|
| latest update_time in the batch | 2024-01-02 11:00:00 |
| update_time of the published value | 2024-01-02 09:00:00 |
Current collapse step (excerpt)
# keep one change per customer before merging into main_data latest = cdc_data.dropDuplicates(['customer_id'])
Investigation task
Apply the CDC batch to main_data so each customer_id ends up with exactly one row carrying its most recent change. Build df_result with customer_id, name and age, and finish with df_result.show().
The fix is written and graded in the regular challenge editor. The postmortem unlocks once the incident is resolved.