I keep hearing the same refrain: “We can’t use AI until we sort our data out.” It sounds sensible. Responsible, even. But I’ve long suspected it doesn’t really hold water.
Having now used AI models and agents to deliver real value on demonstrably dodgy data, I’m convinced that waiting for clean data is the wrong approach. For years, data cleansing has been treated as a sacred prerequisite: before you can build models, deploy agents, or trust insights, you’re told to normalise, deduplicate, reconcile, and perfect your datasets.
The advice is well-intentioned, but increasingly wrong. Not because data quality doesn’t matter, but because “cleaning the data” has become a convenient proxy for avoiding harder questions about purpose, risk, and how to unlock value with the tools we now have.
This is not an argument for ignoring data quality. It’s an argument against treating ‘perfect data’ as a prerequisite for mining value from what you have now.
The uncomfortable truth: most corporate data is never clean, and never will be
I’ve lost count of the data cleansing initiatives I’ve seen (and been part of) over the years. In my experience, they usually cost far more than the benefit they deliver. They soak up huge amounts of analysis, meeting time, and emotional energy!
The reality is that corporate data is sprawling, dynamic, and tied into business processes that are hard to change. It’s owned by different teams, managed in different ways, and constantly evolving. Expecting it to be pristine is unrealistic.
That doesn’t mean bad data should be ignored. Diligence around data quality and integrity absolutely matters. Just don’t expect, or wait for perfection. Many of the most successful AI systems in use today operate on data that is clearly imperfect.
AI is statistical, which matters
Modern AI systems, particularly large language models, aren’t rigid rule engines. They’re statistical learners. That means they can tolerate noise, smooth over anomalies, and extract value from messy, real-world data.
In practice, perfect data is rarely required to get useful results. AI can often help normalise inputs and outputs, highlight inconsistencies, and make sense of ambiguity — even when the underlying data isn’t pretty.
There are, of course, domains such as regulatory reporting, safety-critical systems, financial controls, where higher data standards are non-negotiable. But even there, AI can be used to surface risk, not postponed until risk disappears.
So insisting on pristine datasets before even exploring how AI might deliver value is a cardinal error (pardon the data-modelling pun).
Get more value from the data you already have
I’ve always found that optimism produces better solutions. If you focus first on what might be possible, rather than everything that could go wrong, you’re far more likely to create something useful.
Ask practical questions: What decisions could be influenced with the data that’s already available? Where does uncertainty matter — and where does it not?
If you have a strategy for handling ambiguity and errors, AI can actually help improve data quality over time. As confidence grows and the data improves, you can afford to give AI a longer leash.
Three top tips
Instead of cleaning everything upfront, flip the model:
Start with a real use case and be ambitious
Be clear about the decision you want to improve and the value you’re aiming to unlock.
Deploy AI with humans in the loop
Pay attention to where AI surfaces uncertainty, contradictions, or anomalies — those are signals, not failures.
Fix errors as they’re found, at the source
Use AI as an accelerator to identify and correct issues where they originate.
In this model, data quality improves because the data is being used, not because someone declared a cleansing project.
To find out more about to get more from AI, contact us and let's have a chat.
JOHN PRIDEAUX