The more impressive AI becomes at the things it can do, the easier it becomes to overestimate the things it can reliably do.
Ask an AI to read a clock.
ClockBench does exactly that. It contains 180 deliberately awkward analogue clocks: Roman numerals, mirrored faces, missing numbers, unfamiliar layouts — the sort that reward genuinely careful looking. Humans find the task relatively easy. AI, until recently, did not.
When the benchmark was published in 2025, five adult participants averaged 89.1% accuracy on valid clocks; the best of 11 AI models managed just 13.3%. A year later, the live benchmark contains 36 model entries and the best scores 66.7%, compared with a human baseline of 90.7%.
This is not a clean before-and-after comparison. The human sample was tiny, the models have changed and the live leaderboard keeps changing. But the direction is hard to miss. Something AI found extraordinarily difficult a year ago is becoming much easier.
The errors are perhaps even more revealing. When the humans in the original study got the time wrong, their median error was three minutes. For the best-performing AI model, it was an hour.
The point is not that AI is stupid because it struggles to tell the time. It is that capability is jagged. The same generation of technology can solve difficult mathematical problems, write sophisticated software and reason through specialist material, yet still perform surprisingly badly on another apparently simpler task. Being extraordinarily good at one thing tells us much less than we might expect about performance on another.
That matters beyond benchmarks because organisations can make essentially the same reasoning error twice. First, we infer dependable performance from an impressive demonstration of technical capability. Then, once the technology really is dependable, we infer that enterprise value should follow.
Neither inference is safe.
There is no single thing called “AI”
We routinely say that “AI can now do X” or that “AI has reached human performance”, as though one intelligence were progressing steadily towards one human benchmark. It isn’t.
Human benchmarks usually take a group of people and report some kind of average or baseline. AI leaderboards compare competing models, versions and configurations, and our attention naturally goes to whichever one is at the top. The best AI is no more representative of all AI than the best human is representative of all humans.
ClockBench makes this visible too. Its current results range from 66.7% at the top to below 4% at the bottom. Change the task and the ranking changes with it. Different models have different strengths, while size, modality, reasoning approach, cost and configuration can all matter.
The frontier also moves quickly enough that the tests themselves have to move. When OSWorld was introduced in 2024 to test whether AI agents could operate real desktop applications, humans completed more than 72% of its tasks while the best AI system managed just 12.24%. In 2026 the researchers introduced OSWorld 2.0, built around much longer and more realistic workflows; its best reported system completed only 20.6% under the paper’s main measure.
That is not a fair before-and-after comparison because the second test is deliberately much harder. But that is exactly the point. As systems get better at one class of task, harder tests expose the next weakness.
So “How capable is AI now?” is too broad to be very useful. Capable at what? Which model? Configured how? At what cost? And, for an organisation contemplating real delegation, at what level of reliability?
And even the model is not the whole story.
ARC-AGI-3 asks AI systems to explore unfamiliar environments, infer how they work and achieve goals without being given the underlying rules. GPT-6 Astra achieved 62.7% using ARC Prize’s standard interface. Using a provider-specific setup that preserves the model’s reasoning state between requests and manages longer conversations differently, the best observed result reached 99.9%. ARC Prize treats these as different evaluation questions because the interface matters.
The practical point is simple: the model is not the system. Memory, context, information, tools, retrieval, deterministic logic, validation, workflow and human escalation can materially change what an AI solution can do.
“AI can do this” is therefore an incomplete statement. What matters is whether the system you actually plan to use can do the task reliably enough, at acceptable cost and risk.
So the first gap is simple: demonstrated model capability is not the same thing as dependable operating capability.
But there is another.
From dependable capability to enterprise value
So assume we now have capable technology: the model can perform the task, the surrounding system is reliable enough, and the cost and risk are acceptable.
What does it take to turn that capability into value?
Sometimes, surprisingly little. A large field study of 5,179 customer-support agents found that access to an AI assistant increased issues resolved per hour by 14% on average, with substantially larger gains among less experienced workers. The AI fitted into an existing workflow and helped people do work they were already doing more effectively. In this case, the organisation did not need to redesign the work before seeing a measurable benefit.
That matters because it tells us that organisational change is not automatically the next problem. If an existing workflow can absorb a new capability, value may appear simply because the people doing the work now have a better tool.
But not every opportunity looks like that. A six-month randomised field experiment across 66 firms and 7,137 knowledge workers found bigger changes in behaviours people could adjust for themselves and much less movement where change required coordination with colleagues. That suggests an important distinction: some value can be captured by changing what one person can do, while other value depends on changing what several people, teams or functions do together.
Once AI is intended not just to improve an existing task but to remove activities, redesign an end-to-end process, move a decision from one role to another or automate work surrounded by controls and accountability, the technology is only part of the change. The capability may exist, but realising the value may still require changes to workflow, authority, incentives, controls, responsibilities or coordination.
The difficulty begins when those changes span organisational boundaries. The cost may sit in one function while the benefit appears in another; decision authority may sit somewhere different again; and accountability may remain attached to the old process. Each part of the organisation can therefore behave reasonably within its own responsibilities while nobody has both the authority and the incentive to make the end-to-end change happen.
At that point, improving the model again may achieve remarkably little.
The technology may be good enough. The constraint has moved.
Find the constraint that matters
The debate about whether AI transformation is “a technology problem” or “an organisational problem” is a natural distraction, born from our preference for simple answers. It asks us to choose a general explanation when the real constraint depends on the particular problem in front of us.
Think about a production line. If one stage can process 100 units an hour and the next can process only 50, increasing the capacity of the first stage to 200 does not increase the output of the line. You have improved one part of the system, perhaps substantially, but you have not improved the outcome because something else is limiting it.
AI transformation can work in much the same way. If the technology cannot perform the required task reliably enough, better processes, clearer accountability and enthusiastic adoption will not rescue the proposition. The technology is still the constraint. But once the technology becomes good enough, continuing to improve it may create very little additional value if the organisation still depends on old workflows, approvals, controls or roles that prevent the capability being used differently.
At that point, further improvement in the model is like speeding up the wrong part of the production line.
There is a simple test. Assume the AI technology works perfectly tomorrow.
What would still have to change before the organisation realised the value?
If the answer involves workflow, authority, incentives, funding, controls, accountability or coordination, then implementing the technology will not realise the value on its own. The technical capability may be necessary, but it is no longer sufficient.
The converse is simpler: if nothing about the organisation is preventing the change, ask whether the technology can actually perform the required function reliably enough. If it cannot, calling the problem “change management” merely disguises an inadequate technical proposition.
The point is not to prove that organisation matters more than technology. It is to identify what is preventing the next increment of value now. That constraint can change over time: early in an opportunity the model may simply not be capable enough; later the challenge may be data, integration or engineering the surrounding system; once those conditions are met, ownership, coordination, authority or adoption may become decisive. In another use case, none of those organisational changes may be necessary at all.
So “AI readiness” in the abstract tells us surprisingly little. The more useful question is what currently sits between this particular technical possibility and the value the organisation wants from it — and whether we are investing in that constraint, or simply continuing to improve a part of the system that is already good enough.
The AI Transformation Paradox
AI is improving quickly, but not evenly. There is no single AI moving smoothly towards one measure of human capability. Different models and systems are good at different things, and what looks possible changes with the task, the system around the model and the reliability we require.
But once the technical system works, we make the same mistake again.
We see a model perform an extraordinary task and infer that AI can reliably do the job. Then we build a system that can reliably do the job and infer that the value should follow.
The paradox is that each improvement in capability can make the next constraint easier to overlook.
The faster the technology improves, the more important it becomes to distinguish what is possible from what is dependable, and what is dependable from what will actually create enterprise value.
So the question is no longer simply:
What can AI do?
It is:
What can this system do reliably enough to matter — and what is stopping us capturing the value if it can?