Revenue Inflated After a Customer Feed Load
PySpark · data quality · Intermediate · about 20 minutes
Impact
The revenue dashboard over-reports sales for some customers after last night's customer-feed load, while order volume has not changed.
Symptoms
- The ETL job finished with status SUCCESS and raised no errors.
- Revenue for customer 1 is exactly double what the orders table says.
- Revenue for customer 2 is correct.
- The joined orders table has more rows than the orders it was built from.
Evidence
Row counts, last run
| dataset | rows |
|---|---|
| df_orders (input) | 3 |
| df_customers (lookup feed) | 3 |
| orders joined to customers (output) | 5 |
Revenue check, customer 1
| source | revenue |
|---|---|
| Sum of df_orders.amount | 340 |
| Revenue dashboard | 680 |
df_customers, as delivered by the feed
| customer_id | name |
|---|---|
| 1 | Alice |
| 1 | Alice_dup |
| 2 | Bob |
Job log
[etl] load_customers rows_written=3 status=OK [etl] join_orders_customers rows_in=3 rows_out=5 status=OK [etl] publish_revenue status=SUCCESS
Investigation task
Rebuild the orders-to-customers join so every order appears exactly once with one customer name. Build a DataFrame called df_result and finish with df_result.show().
The fix is written and graded in the regular challenge editor. The postmortem unlocks once the incident is resolved.