Home / Expertise / AI data readiness

AI data readiness

Test whether the use case can be evaluated before selecting a model.

AI readiness begins with a decision, a measurable outcome and a credible feedback path. Data volume alone cannot establish feasibility, safety or value.

The actual problem

The proposed model has a name, but the decision and ground truth do not.

Many AI initiatives begin with a technology category and search for a use case afterwards. This postpones the hardest questions: what action will change, how a correct outcome will be observed, which errors matter and whether historical data represents future operating conditions.

A readiness assessment tests a specific decision pathway. It can conclude that the use case is ready for a bounded experiment, requires evidence work first, should use a simpler rule or analysis, or should not proceed.

  • Technology-first scopeThe initiative is described as a model or assistant without a named operational decision.
  • Unavailable ground truthThe desired outcome is subjective, delayed, inconsistently recorded or affected by the intervention itself.
  • Historical mismatchPast cases reflect policies, populations or workflows that will differ at deployment.
  • Accuracy without baselineA target metric is proposed without the class rate, current process or cost of each error.
  • Feature leakageTraining data contains information that would not exist when a live prediction is required.
  • No feedback ownerNobody is accountable for reviewing drift, overrides, harm or whether the recommendation improved the outcome.

Method

Assess the complete decision system, not only the training dataset.

A technically feasible prediction can still be operationally useless or unsafe. Readiness includes data, evaluation, workflow, accountability and the ability to learn from outcomes after deployment.

Decision and action
Define who receives the output, what action it changes, when it is needed and where human judgement or policy remains authoritative.
Outcome and ground truth
Specify the target, observation window, label process, ambiguity, delay and how intervention changes the observed outcome.
Population and representation
Compare historical and intended populations across time, segment, geography, channel and relevant operating conditions.
Feature availability
Confirm that candidate inputs exist at decision time, have stable meaning and can be produced with acceptable latency and quality.
Evaluation and error cost
Select metrics against base rates, current practice, segment performance and the operational consequence of false positives and false negatives.
Operation and feedback
Design review, override, monitoring, incident response, drift detection, outcome capture and retirement conditions.

Evidence required

What must be known before a credible experiment can be designed.

The assessment requires representative records and the operating process around them. A dataset extracted without policy history, timestamps or label provenance can make an infeasible use case look ready.

InputReadiness questionMaterial risk
Decision workflow and current baselineWhat action changes and how well does the current process perform?Optimising a prediction that cannot improve the actual workflow.
Historical cases with event timeDo records represent what was known at the decision moment?Leakage from later information or retrospective correction.
Label definition and provenanceWho assigned the outcome, under which policy and with what consistency?Learning inconsistency, bias or a proxy unrelated to value.
Population and segment profileWill deployment encounter the same distribution and edge cases?Strong aggregate results hiding weak or harmful segment performance.
Error and intervention costsWhat happens when the system is wrong or ignored?An attractive metric paired with an unacceptable operating consequence.
Monitoring and feedback capabilityCan outcomes, overrides, drift and incidents be observed after launch?Performance degradation without detection or accountable response.

Outputs

A readiness decision with explicit evidence gaps and safeguards.

The output does not promise model performance. It states whether a responsible evaluation is possible and what must be true before operational use.

  • Use-case decision contractUser, decision, action, timing, constraints, human authority and intended outcome.
  • Data and label assessmentCoverage, provenance, quality, representativeness, leakage risks and ground-truth limitations.
  • Baseline and evaluation designCurrent process, base rates, time-aware split, segment checks and error-cost measures.
  • Readiness dispositionProceed to bounded experiment, complete evidence work, use a simpler approach or stop.
  • Experiment specificationPopulation, comparison, success criteria, guardrails, review authority and stopping rules.
  • Operational control planMonitoring, overrides, feedback capture, incident response, drift review and retirement conditions.

Worked example

Ninety-seven percent accuracy can represent no useful model at all.

A team proposes a model to flag cases likely to require escalation. There are 10,000 historical cases, but an outcome label is present for only 6,200. Among labelled cases, 3% ended in escalation.

Readiness elementObserved valueImplication
Label coverage6,200 / 10,00038% of cases have no evaluable outcome.
Positive base rate3% of labelled casesApproximately 186 positive examples exist.
Naive baselinePredict no escalationAchieves 97% accuracy while finding zero escalations.
25% validation shareApproximately 1,550 casesExpected positive cases are only about 47 before segment analysis.
Label coverage = 6,200 / 10,000 = 62%
Positive examples = 6,200 x 3% = 186
Naive no-escalation accuracy = 1 - 3% = 97%
Expected validation positives = 6,200 x 25% x 3% = 46.5

Accuracy is unsuitable as the primary success measure. Readiness depends on why labels are missing, whether historical policy influenced escalation, the relative cost of missed and unnecessary interventions, and whether there are enough positive examples to evaluate relevant segments over time.

The correct next step is label and workflow diagnosis, followed by a time-aware evaluation against the current triage process using precision, recall, intervention capacity and error cost, not a broad claim of model readiness.

Limits

Readiness is not a performance guarantee or safety certification.

The assessment establishes whether a use case can be evaluated responsibly with available evidence and controls. Actual performance emerges only through bounded testing and monitored operation.

  • Not model selectionAlgorithm and vendor choices follow the decision, evidence and evaluation design.
  • Not proof of causal valuePredictive performance does not establish that acting on a prediction improves the outcome.
  • Not a complete legal or ethical reviewApplicable rights, obligations, prohibited uses and broader social effects require qualified assessment.
  • Not permanent approvalPopulation, policy, data and behaviour change; readiness and controls must be reviewed over time.

Which proposed AI use case needs an evidence test?

Assess readiness