Home / Expertise / Data quality review

Data quality review

Determine whether the data is fit for the decision.

A quality review links defects to their effect on metrics and decisions. It distinguishes inconvenient data from evidence that is materially unreliable.

The actual problem

Trust declines faster than anyone can explain the defect.

“The data is bad” is not a useful diagnosis. Quality is contextual: a two-day delay may be harmless for quarterly planning and unacceptable for daily capacity management. A missing attribute may leave a total intact while making segment decisions impossible.

The review starts from the decision and metric, traces the required data and measures defects against explicit expectations. This produces a prioritised risk view rather than a catalogue of every irregularity in the estate.

  • Manual reconciliationTeams repeatedly adjust extracts or maintain shadow spreadsheets before using a report.
  • Unexplained restatementsHistorical values change after refreshes without a visible correction policy.
  • Broken segmentationTotals appear plausible, but missing or inconsistent dimensions prevent meaningful comparison.
  • Late-data surprisesA period is discussed before material records have arrived or settled.
  • Untraceable metricsNo one can reproduce a headline value from source records and transformation logic.
  • Alert fatigueHundreds of technical tests fail without indicating which business outcome is at risk.

Method

Assess quality along the path from source to decision.

The method combines data profiling with semantic and operational review. Passing a format test is not enough if the records represent the wrong population or arrive after the decision is made.

Critical-data scope
Select the metrics, decisions and source elements whose failure could materially change an action or external statement.
Expectation definition
Set rules for completeness, uniqueness, validity, consistency, timeliness and reconciliation at the relevant grain.
Source profiling
Measure distributions, missingness, duplicates, invalid values, drift and arrival patterns over representative periods.
Transformation trace
Follow filters, joins, aggregations, backfills and late-arriving records from source to presented metric.
Impact assessment
Quantify how each defect changes totals, rates, segments, rankings or the confidence of the decision.
Control design
Place detection and response at the earliest practical point, with an owner, tolerance, escalation and correction rule.

Evidence required

A review needs records, logic and operating context.

Documentation alone cannot establish actual quality, while profiling alone cannot establish whether a defect matters. Both are required.

InputPurposeReview question
Representative raw extractsProfile actual values, keys, timestamps and distributions.Does the sample cover normal, peak and known incident periods?
Data model and grainEstablish what one row represents and which keys should be unique.Can joins multiply or suppress business entities?
Transformation logicTrace filters, derivations, aggregations and restatements.Where can a valid source record disappear or change meaning?
Metric definitionsConnect technical defects to reported values.Which fields affect numerator, denominator or segmentation?
Arrival and incident historyMeasure freshness, recovery and recurring failure patterns.When is a period sufficiently complete for use?
Decision calendarSet business tolerances for latency and correction.What happens if the issue is discovered after the decision?

Outputs

Evidence of fitness, not a generic quality score.

A single percentage hides which defects matter and which decisions remain supportable. The output preserves that distinction.

  • Critical data inventoryMetrics and decisions mapped to source elements, transformations, owners and required service levels.
  • Quality profileMeasured completeness, uniqueness, validity, consistency, timeliness and distribution stability.
  • Impact analysisQuantified effect of defects on totals, rates, segments, trends and decision confidence.
  • Risk-ranked findingsIssues ordered by business materiality, recurrence, detectability and effort to correct.
  • Control specificationRule, tolerance, location, owner, alert route and expected response for each priority risk.
  • Correction and restatement policyHow late or corrected data changes published results and how users are informed.

Worked example

A plausible total can still overstate the reporting population.

A monthly customer extract contains 100,000 rows and is used as the population for an activity rate. Profiling finds duplicate account-month keys, missing plan tiers and records arriving after the reporting cutoff.

FindingRows affectedEffect
Duplicate account-month rows4,200 excess rowsInflates the population and any additive totals.
Missing plan tier8,100 rowsTotal remains available; plan-level comparison is incomplete.
Arrived after reporting cutoff1,700 rows, including 350 duplicatesCreates a mismatch between as-of and later-restated views.
Unique account-month population = 100,000 - 4,200 = 95,800
Unique rows available by cutoff = 95,800 - (1,700 - 350) = 94,450
Original overstatement versus as-of population = (100,000 - 94,450) / 94,450 = 5.88%

The missing plan tier does not invalidate the overall account count, but it makes plan-level rates unsafe unless the missingness is understood and disclosed. Duplicates and late records directly affect the as-of denominator.

The control design should therefore separate uniqueness, segment completeness and cutoff freshness rather than compressing them into one “data quality” indicator.

Limits

A bounded review is not a certification of the data estate.

The findings apply to the scoped sources, periods, transformations and decisions. Unknown defects can remain outside that boundary.

  • Sampling has boundariesRepresentative periods reduce risk but cannot prove that every historical or future record is correct.
  • Validity needs business meaningA technically valid value may still describe the wrong entity, event or period.
  • Detection is not remediationSource-process, pipeline and ownership changes require implementation by the responsible teams.
  • Quality is decision-specificPassing one use case does not make the same data suitable for another purpose with tighter tolerances.

Which decision is most exposed to unreliable data?

Scope a quality review