Case Study: Taming Mixed-Format Supplier Data
Updated
A fictional retailer receives product feeds as CSV, JSON and one XML 'because the supplier is special' - the project is a validation pipeline with a data-loss report a buyer can actually read.
Fictional training case, not a real engagement.
Situation
"Harbor & Lane" (fictional) onboards supplier product feeds into their catalog. Each supplier has quirks: encodings (windows-1254!), unit formats ("2x500ML"), duplicate GTINs, price format drift. Manual fixes cost ~9 h/week.
The delivery arc
- Discovery: 11 suppliers; 4 formats; current process is "the intern fixes it"; error rate unknown - first deliverable is a measurement harness.
- Metrics: manual fix time 9h → <2h weekly (60 days); zero silent drops (hard); quarantine report read by a non-technical buyer.
- Pipeline: format adapters → canonical model → validation rules (nullable price? VAT field? unit parsing) → quarantine with human-readable reasons → catalog load.
- The interesting failure: one supplier's "special" XML shipped malformed CDATA twice a month - the pipeline's quarantine saved the run, and the report convinced the supplier to fix their export.
- Handover: rule documentation written for the ops analyst who now maintains thresholds.
What learners should extract
- Measurement harnesses are legitimate first deliverables (why baselines matter).
- Quarantine + readable error taxonomy beats silent best-effort loading.
- "Zero silent drops" is the trust metric buyers understand.
Practice version
Our data validation practice task is a mini version with three deliberately broken rows per format.