An AI-ready data strategy is one that will deliver a environment where your data is consistent enough, has clearly named owners, has a core 'golden source' for each data set (sales, customer, cost etc) and accessible enough that an AI project can start without a six-month clean-up first.
For most UK SMEs the blocker is not model access. It is that data is not connected, key metrics are not clearly not defined and two systems disagree about what a customer or sale actually means.
The pattern repeats. A business decides to do something with AI, scopes a use case, and then discovers that the data it needs sits across three systems that are not connected and were never designed to agree. The project does not fail suddenly but it quietly becomes a data integration project with an AI name on it, the original sponsor loses interest and users don't trust the outputs enough to actually use them.
This is not an argument for delaying AI work. It is an argument for knowing which of the two projects you are actually funding before you start. If you start with AI before data then the project is considerably more likely to never get off the ground.
Four things, in roughly this order of difficulty:
| Requirement | What it means in practice | What happens without it |
|---|---|---|
| Agreed definitions | One answer to what counts as a customer, an order, an active account, coming out of core system that reports on it | The model produces confident answers to a question you never actually defined |
| Traceable lineage | For any number, you can say where it came from and what happened on the way. | Nobody can debug an output that is claimed to 'look wrong', so nobody trusts any output |
| Named ownership | A person, not a team, owns each core definition and decides when it changes | Definitions drift silently and the previous quarter is no longer comparable to the latest one |
| Accessible history | Enough clean historical data to establish what normal looks like | There is no past worth learning from, so prediction is guesswork with a confidence score |
Notice that none of these are AI work. They are the foundations that make AI work possible, and they are worth doing whether or not you ever train a model, because they also fix your reporting. That is the whole basis of our data foundations engagement.
Pick the three to five key numbers the business actually runs on. Trace each one to source. Where recorded in more than one system, document where they disagree on how it is defined, and decide which one wins. This stage produces arguments rather than deliverables, which is why it gets skipped, and why skipping it is expensive.
Move those definitions into a tool that can be automated, so the numbers are produced the same way every time without someone rebuilding a spreadsheet. At this point reporting becomes trustworthy, which is usually the first moment the business feels a return. Our analytics platform work covers this layer.
Choose a single decision that would change if you could predict something. Not the most impressive use case but the one that matters in the business and currently leads to the most discussion or disagreement. Build it, measure whether the decision improved, and be willing to stop if it did not.
Only after one use case has demonstrably worked. If you widening to further use cases before that point it will simply multiply an unproven approach across the business, which is how organisations end up with several half-finished AI projects and confidence in what they can deliver disappears.
Three failures account for most of the issues.
(1) Buying tooling before agreeing definitions, so the tool faithfully reports the disagreement.
(2) Choosing a use case nobody has authority to act on, so a correct prediction changes nothing.
(3) Treating readiness as a one off activity, rather than something that becomes repeatable as systems and teams change.
If your AI work touches personal data, the ICO's guidance on AI and data protection sets out what you need to be able to demonstrate. Worth reading at stage one of your road map rather than stage four, because it shapes what you are allowed to build.
The same key requirements apply, but the ownership question gets harder, because the named owner of a definition has to sit with someone in the business. In practice that usually means a finance or operations lead taking responsibility for the numbers in their area, with external help doing the engineering.
This will be driven by the extent to which your are systems and data are currently connected and key output metrics defined, not the size of the business. Reconciling a handful of core definitions across a few systems is a fairly short exercise. If every department maintains its own version of the truth, expect the conversation to take longer than the engineering.
Not necessarily, but it helps. Without one you will need somewhere the agreed version of the data lives that is not a spreadsheet on someone's machine. Whether that is a warehouse, a lakehouse or a simpler database depends on volume and how many systems need to feed it. Regardless of the option you choose, the ingestion of data from your core systems needs to be automated and repeatable without manual intervention.
The fundamentals requirements of fixing reporting and becoming AI-ready are the same, connected data, clear definitions and accountable owners. If you have to choose, choose the one that delivers the highest pay back for where you are now.
If you are working out which stage you are at, our data foundations work covers stages one and two, and advisory support covers the sequencing decisions above. Our case studies show what the first two stages produced for a business with the same problem.