FDE Foundations · Module 5: Data & Integration
Data Quality Checks
Data quality is four checks (complete, unique, valid, fresh) with thresholds someone agreed to, reported on a schedule the customer actually reads.
10 min reading
Objectives
- Define completeness, uniqueness, validity, and freshness checks
- Pick thresholds that trigger action, not noise
- Publish a weekly data quality report
The four checks
Completeness: every load delivered the expected rows, and required fields are populated. Uniqueness: no duplicate business keys after idempotency logic. Validity: values pass format and range rules; foreign identifiers resolve. Freshness: the newest record is within the agreed window. Each check is cheap; together they catch most field incidents before users do.
Thresholds need an owner
A check without a threshold and an owner is decoration. Agree numbers: "duplicate rate above 0.5 percent pages the on-call", "freshness older than 3 hours opens a ticket". Vague checks train everyone to ignore them.
The report
Publish a weekly one-pager: rows in, rows quarantined, top three failure reasons, trend arrows. This report does double duty as the engagement's honesty artifact: it shows the customer problems you found before they did, which is exactly what an FDE is for.
Fix sources, not symptoms
When a check fails repeatedly, walk the failure upstream. The durable fix is often a schema change at the source system or a validation rule in the sender's export job, not a smarter parser on your side.
Quick check
An optional 2-3 question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Exercise
Write the data quality plan for a fictional customer-master sync: one check per category (complete, unique, valid, fresh) with threshold, action, and owner, plus the top line of the weekly report.
Pass criteria
Four checks each with a numeric threshold, a named action, and an owner; the report top line states volumes and the leading failure reason.