Case Study: Taming Mixed-Format Supplier Data

Updated

A fictional retailer receives product feeds as CSV, JSON and one XML 'because the supplier is special' - the project is a validation pipeline with a data-loss report a buyer can actually read.

Fictional training case, not a real engagement.

Situation

"Harbor & Lane" (fictional) onboards supplier product feeds into their catalog. Each supplier has quirks: encodings (windows-1254!), unit formats ("2x500ML"), duplicate GTINs, price format drift. Manual fixes cost ~9 h/week.

The delivery arc

  1. Discovery: 11 suppliers; 4 formats; current process is "the intern fixes it"; error rate unknown - first deliverable is a measurement harness.
  2. Metrics: manual fix time 9h → <2h weekly (60 days); zero silent drops (hard); quarantine report read by a non-technical buyer.
  3. Pipeline: format adapters → canonical model → validation rules (nullable price? VAT field? unit parsing) → quarantine with human-readable reasons → catalog load.
  4. The interesting failure: one supplier's "special" XML shipped malformed CDATA twice a month - the pipeline's quarantine saved the run, and the report convinced the supplier to fix their export.
  5. Handover: rule documentation written for the ops analyst who now maintains thresholds.

What learners should extract

  • Measurement harnesses are legitimate first deliverables (why baselines matter).
  • Quarantine + readable error taxonomy beats silent best-effort loading.
  • "Zero silent drops" is the trust metric buyers understand.

Practice version

Our data validation practice task is a mini version with three deliberately broken rows per format.

We use Google Analytics to count visits. No ads, no cross-site tracking. Cookie Policy