What is Data Readiness?
Strategy & BuyingWhether the data an AI system needs actually exists, can be reached, is good enough to learn from, and may lawfully be used. Six questions asked before the build, because the answers decide what can be built at all.
Why It Matters
Every AI project begins with a claim about data, and the claim is usually optimistic. “We have tons of data” describes volume and says nothing about whether it is labelled, current, legally usable, or relevant to the question being asked. Discovering that six weeks into a build is the most common way an AI project becomes a data project without anyone agreeing to it.
Data readiness is the practice of asking the questions before the build rather than during it. It is cheap as an assessment and expensive as a discovery, which is why it belongs in a feasibility study rather than in the first sprint.
The Six Dimensions
Availability. Do we have the data at all, can we reach it, and can we extract it without a project of its own? Data that exists inside a system nobody can query is not yet available.
Quality. Is it accurate, complete, and consistent, and is it up to date? Poor data produces poor output, and no amount of model tuning repairs a broken pipeline.
Volume. Is there enough of it, is it representative of the real world, and does it cover the edge cases? A model trained on the easy cases will be confident about the hard ones.
Relevance. Does it relate to the problem, does it contain the signals the decision needs, and is it labelled where labelling is required? Relevant data for one question is noise for another.
Security. Is it properly protected, who has access, does it meet the regulations that apply, and can it be anonymised where it must be? The access question usually reveals that the answer is “more people than expected”.
Governance. Do we have permission to use it for this purpose, is there a data owner, and is the retention defined? This is the dimension that stops a technically perfect project, and it is the cheapest to check first.
How to Use It
Score each dimension honestly, then let the weakest one set the scope. If governance fails, nothing else matters until it is resolved. If volume is thin, the design changes to retrieval over documents rather than training on examples. If quality is poor, the first deliverable is a data pipeline, not an agent.
The output is not a scorecard. It is a decision about what to build first, and an explicit list of what was ruled out and why.
Where It Breaks
Volume mistaken for readiness. The most common error, and the one the phrase “we have tons of data” encodes. Volume is one of six questions and rarely the binding constraint.
Assessing data instead of access. A catalogue that lists the data is not the same as a system that returns it. The assessment has to touch the real source, not a metadata inventory.
Treating governance as paperwork. Permission, ownership, and retention are design inputs. Discovering them late means rebuilding the scope rather than adding a document.
A readiness score with no decision attached. A number that does not change what gets built has measured nothing that matters.
Assuming retrieval avoids the question. Working over documents rather than training on examples lowers the bar on volume and labelling. It does not lower it on availability, relevance, security, or governance.
How Flytebit Handles It
The six dimensions are part of every feasibility study, scored against the real sources rather than a data catalogue, and the weakest one sets the first build. Where documents are the material, the work continues into retrieval rather than training, which changes which dimensions bind. The assessment sits inside the wider readiness work that decides whether an AI project should start at all.
More info
- EU AI Act: consolidated text (data and data governance) Data governance is a named requirement for high-risk systems, not a preparatory step.