AI data readiness
Test whether the use case can be evaluated before selecting a model.
AI readiness begins with a decision, a measurable outcome and a credible feedback path. Data volume alone cannot establish feasibility, safety or value.
The actual problem
The proposed model has a name, but the decision and ground truth do not.
Many AI initiatives begin with a technology category and search for a use case afterwards. This postpones the hardest questions: what action will change, how a correct outcome will be observed, which errors matter and whether historical data represents future operating conditions.
A readiness assessment tests a specific decision pathway. It can conclude that the use case is ready for a bounded experiment, requires evidence work first, should use a simpler rule or analysis, or should not proceed.
- Technology-first scopeThe initiative is described as a model or assistant without a named operational decision.
- Unavailable ground truthThe desired outcome is subjective, delayed, inconsistently recorded or affected by the intervention itself.
- Historical mismatchPast cases reflect policies, populations or workflows that will differ at deployment.
- Accuracy without baselineA target metric is proposed without the class rate, current process or cost of each error.
- Feature leakageTraining data contains information that would not exist when a live prediction is required.
- No feedback ownerNobody is accountable for reviewing drift, overrides, harm or whether the recommendation improved the outcome.
Method
Assess the complete decision system, not only the training dataset.
A technically feasible prediction can still be operationally useless or unsafe. Readiness includes data, evaluation, workflow, accountability and the ability to learn from outcomes after deployment.
- Decision and action
- Define who receives the output, what action it changes, when it is needed and where human judgement or policy remains authoritative.
- Outcome and ground truth
- Specify the target, observation window, label process, ambiguity, delay and how intervention changes the observed outcome.
- Population and representation
- Compare historical and intended populations across time, segment, geography, channel and relevant operating conditions.
- Feature availability
- Confirm that candidate inputs exist at decision time, have stable meaning and can be produced with acceptable latency and quality.
- Evaluation and error cost
- Select metrics against base rates, current practice, segment performance and the operational consequence of false positives and false negatives.
- Operation and feedback
- Design review, override, monitoring, incident response, drift detection, outcome capture and retirement conditions.
Evidence required
What must be known before a credible experiment can be designed.
The assessment requires representative records and the operating process around them. A dataset extracted without policy history, timestamps or label provenance can make an infeasible use case look ready.
| Input | Readiness question | Material risk |
|---|---|---|
| Decision workflow and current baseline | What action changes and how well does the current process perform? | Optimising a prediction that cannot improve the actual workflow. |
| Historical cases with event time | Do records represent what was known at the decision moment? | Leakage from later information or retrospective correction. |
| Label definition and provenance | Who assigned the outcome, under which policy and with what consistency? | Learning inconsistency, bias or a proxy unrelated to value. |
| Population and segment profile | Will deployment encounter the same distribution and edge cases? | Strong aggregate results hiding weak or harmful segment performance. |
| Error and intervention costs | What happens when the system is wrong or ignored? | An attractive metric paired with an unacceptable operating consequence. |
| Monitoring and feedback capability | Can outcomes, overrides, drift and incidents be observed after launch? | Performance degradation without detection or accountable response. |
Outputs
A readiness decision with explicit evidence gaps and safeguards.
The output does not promise model performance. It states whether a responsible evaluation is possible and what must be true before operational use.
- Use-case decision contractUser, decision, action, timing, constraints, human authority and intended outcome.
- Data and label assessmentCoverage, provenance, quality, representativeness, leakage risks and ground-truth limitations.
- Baseline and evaluation designCurrent process, base rates, time-aware split, segment checks and error-cost measures.
- Readiness dispositionProceed to bounded experiment, complete evidence work, use a simpler approach or stop.
- Experiment specificationPopulation, comparison, success criteria, guardrails, review authority and stopping rules.
- Operational control planMonitoring, overrides, feedback capture, incident response, drift review and retirement conditions.
Worked example
Ninety-seven percent accuracy can represent no useful model at all.
A team proposes a model to flag cases likely to require escalation. There are 10,000 historical cases, but an outcome label is present for only 6,200. Among labelled cases, 3% ended in escalation.
| Readiness element | Observed value | Implication |
|---|---|---|
| Label coverage | 6,200 / 10,000 | 38% of cases have no evaluable outcome. |
| Positive base rate | 3% of labelled cases | Approximately 186 positive examples exist. |
| Naive baseline | Predict no escalation | Achieves 97% accuracy while finding zero escalations. |
| 25% validation share | Approximately 1,550 cases | Expected positive cases are only about 47 before segment analysis. |
Positive examples = 6,200 x 3% = 186
Naive no-escalation accuracy = 1 - 3% = 97%
Expected validation positives = 6,200 x 25% x 3% = 46.5
Accuracy is unsuitable as the primary success measure. Readiness depends on why labels are missing, whether historical policy influenced escalation, the relative cost of missed and unnecessary interventions, and whether there are enough positive examples to evaluate relevant segments over time.
The correct next step is label and workflow diagnosis, followed by a time-aware evaluation against the current triage process using precision, recall, intervention capacity and error cost, not a broad claim of model readiness.
Limits
Readiness is not a performance guarantee or safety certification.
The assessment establishes whether a use case can be evaluated responsibly with available evidence and controls. Actual performance emerges only through bounded testing and monitored operation.
- Not model selectionAlgorithm and vendor choices follow the decision, evidence and evaluation design.
- Not proof of causal valuePredictive performance does not establish that acting on a prediction improves the outcome.
- Not a complete legal or ethical reviewApplicable rights, obligations, prohibited uses and broader social effects require qualified assessment.
- Not permanent approvalPopulation, policy, data and behaviour change; readiness and controls must be reviewed over time.