Author: IRPA AI Senior Analyst, Kieran Gilmurray

AI use has become ordinary. Most established firms now have it running somewhere, and few executives still ask whether the technology works. Yet more than half of chief executives, 56% in one global survey, report that AI has brought them neither higher revenue nor lower cost over the past year. The obvious reading is that the models are not good enough yet. In a great many organisations, that is the wrong diagnosis.

This article looks at where the ROI constraint actually sits. It argues that once a model is good enough for the job, the factors that decide whether it pays move out of the technology and into the organisation around it.

The Comfortable Diagnosis

When an AI programme fails to pay for itself, the first suspect is almost always the model. The reasoning feels sound: the output was not quite good enough, so the answer must be a better version, a different vendor, more fine-tuning, a longer context window. It is a comfortable diagnosis because it keeps the problem inside the technology, where it can be solved by a purchase rather than a reorganisation. Someone else builds the fix, and the business waits.

The trouble is that swapping the model rarely moves the financial result. In most established enterprises, the model already in hand could do the task to an acceptable standard, and the gain simply never reached the accounts. A more capable model applied to the same unreformed process produces a slightly better demonstration and the same return. The constraint was elsewhere, and buying more capability at the point that was never binding is an expensive way to stand still.

This is not an argument that model quality is irrelevant. It is an argument about where the binding constraint usually sits once a competent model has been chosen. For a growing share of enterprise use cases, the honest answer is that the technology was ready and the organisation was not.

A Model Is a Component, Not a Business Case

A model is a component. It drafts, predicts, classifies and recommends, and it can now act across several steps. None of those outputs is an economic outcome on its own. A recommendation is not a decision, a draft is not a delivered service, and a prediction changes nothing until someone acts on it differently from before. The distance between a capable output and a captured pound is filled entirely by organisational assets, and the model supplies none of them.

Put plainly, a model can generate an answer, but only an operating model can turn that answer into an outcome. The capability creates possibility. Whether the possibility becomes value depends on the workflow it sits in, the data it can reach, the decisions it is allowed to change, the people who use it and the way results are measured. Those are management’s to build, and they are where most of the return is won or lost.

Where the Value Leaks

If the model is rarely the binding constraint, it helps to know what is. The same complements recur across the evidence, and each one can quietly reduce the return of the whole system to very little. Data is the most common. A capable model fed inaccessible, delayed, poorly labelled or unrepresentative data produces less value than a modest model on a reliable information supply. The problem is usually not that data is dirty in the abstract, but that it cannot be retrieved, in the right form, at the moment the workflow needs it.

The workflow itself is the next leak. A model can shorten one step while the process around it still waits on serial approvals, fragmented systems, manual re-entry and queues that never move. Faster contract summarisation changes nothing if legal review stays sequential and sign-off still takes days. The unit of transformation is the workflow or the decision, not the prompt, and a task made quicker inside an unchanged process is a local comfort with no economic weight.

Decisions, and the rights around them, are a subtler leak. An accurate recommendation has little operational value when no one has settled what the system may recommend, what it may decide, what it may execute, who can overrule it and who is accountable for the result. Enterprises repeatedly report that unclear decision and escalation rights are among the strongest brakes on adoption, because a recommendation nobody is authorised to act on is just expensive advice.

Then come adoption, governance and measurement, treated as afterthoughts and nothing of the sort. People will not redesign their work if the incentives, targets and manager behaviour around them still reward the old way, and study after study finds that reinvention is rarely rewarded when short-term numbers slip. Governance decides which actions can safely be delegated to a machine, which is what lets AI move into work worth automating. And measurement, done properly, forces the clarity that makes the rest real: name the business metric that should move, and the vague hope of “productivity” resolves into a case that can be proved or abandoned.

The Adequacy Threshold

It helps to give the underlying idea a name, because a name disciplines the conversation. Every use case has an Adequacy Threshold: the level of capability, reliability, cost and safety at which a model becomes good enough to do the job. Below that line, the model really is the binding constraint, and no amount of redesign will rescue a system that cannot perform the task. At or above it, more model quality buys steeply diminishing returns, and the constraint moves into the organisation. The whole “rarely the model” claim is really a claim about where most enterprise use cases now sit relative to that line.

The threshold also fixes the order of good decisions. First, define what adequate means for this specific use case, in concrete terms: required accuracy, tolerable error, latency, cost per transaction, explainability, and where human judgement is mandatory. Then select the cheapest model that clears the line. Then, and only then, turn the bulk of management attention to the complements that decide the return. Organisations that skip the first step force unsuitable models into valuable workflows. Organisations that never reach the third keep upgrading a model that was already adequate, and keep wondering why the accounts do not move.

The word to hold onto is “rarely,” not “never.” Across a whole portfolio, in an era when strong models are broadly available, the binding constraint is usually organisational. On a single hard use case, complex reasoning, a high-stakes decision, a specialist domain, long-horizon autonomy, the model can still be the thing in the way. The Adequacy Threshold is what tells the two situations apart, and knowing which one you are in is the difference between fixing the right problem and buying the wrong solution.

What It Looks Like in the Work

The pattern is clearest inside specific workflows. Consider an insurer using AI to review claims. The model may extract the facts, classify the claim, flag missing evidence and draft the customer letter, and do all of it well. The return still stalls if the organisation has not decided which claims the system may approve, which need a human, who can overrule it and how exceptions are recorded. The capability was never the issue. The absence of a decision architecture was.

Procurement tells the same story from another angle. A team deploys AI to draft tenders and summarise supplier responses, and the documents appear faster. The sourcing cycle does not shorten, because requirements are still argued out across meetings, evaluation criteria are unclear, legal review runs in sequence, approvals pass through several committees and savings are never tracked after award. The model solved document production. The economic bottleneck was governance and coordination, which no model touches.

Even the commercial model can decide the outcome. In a professional-services firm, a tool that cuts research and drafting time by a wide margin can reduce revenue under hourly billing, improve margin under fixed fees, or do nothing at all, depending entirely on whether pricing, capacity and demand are adjusted around it. The same task-level gain becomes margin, lost revenue or noise according to a business-model choice made far from the technology. This is why “AI made the task faster” is never, on its own, a business case.

None of these failures is a model failure, and none would be cured by a better one. They are failures of the system the model was dropped into: the data it could reach, the decisions it was allowed to change, the process it ran inside, the way value was or was not measured. Replace the model and the problem remains. Redesign the system and a merely adequate model is often enough.

When It Really Is the Model

Honesty requires the other side of the line. There are use cases where model quality is precisely the binding constraint, and pretending otherwise wastes effort in the opposite direction. Complex multistep reasoning, high factual precision, specialist domains, reliable tool use, long-horizon planning and dependable behaviour under unusual conditions can all exceed what today’s models do reliably, and analysts are right to warn that current systems often lack the maturity to pursue complex goals autonomously over time. In those cases a weak model cannot be saved by good change management, and the discipline is to set the capability threshold honestly and refuse to force an inadequate model into a workflow that matters.

The point is symmetry. Organisations should not blame the model for every management failure, and vendors should not blame the organisation for every technical one. The Adequacy Threshold exists to locate the constraint rather than to assume it, so that money and attention go where the binding limit actually is.

The Agent Trap

Agents raise both the ceiling and the stakes. A system that observes, plans and acts across multiple steps can capture far more value than a copilot, but it needs far more of the organisation to be in order: identity, permissions, orchestration, system access, cost control, monitoring, exception handling and clear accountability. Drop an agent into a poor process and it does not repair the process, it executes the waste faster, at higher cost, with errors that surface later and accountability that is harder to trace.

The market is arranging itself for a second pilot trap. Analysts expect a large share of agentic projects to be cancelled within a couple of years, not because the models cannot act, but because costs escalate, value stays unclear and controls are inadequate. They also warn of “agent washing,” where an ordinary assistant or automation is relabelled as an autonomous agent, setting expectations the underlying system was never built to meet. The visible model is the smallest part of the risk.

So agents do not weaken the argument of this piece. They sharpen it. As AI moves from answering to acting, the quality of the surrounding organisation, its data, decision rights, governance and unit economics, becomes more consequential, not less. The firms that capture agentic value will not be the ones with the most autonomous model. They will be the ones whose operating model can be trusted to let a machine act.

What This Means for Leaders

The first move is a diagnosis, not a purchase. Before approving the next model or platform, a leadership team should ask whether it is solving a model problem or an operating-model problem, and be honest about which. If a competent model is already in hand, the productive work is almost never more capability. It is the workflow that has to carry the gain, the data that has to be made reachable, the decisions that have to change and the metric that has to close the loop.

That work needs an owner who owns the outcome, not only the technology. Every material initiative should carry a value hypothesis stated by a business leader: this capability will change this task, relieve this constraint and move this business metric. It needs a baseline, because without one the organisation is left with self-reported time savings and anecdotes. And it needs the value-capture mechanism decided in advance, so that freed capacity becomes throughput, cost, quality or revenue rather than quiet slack or more meetings.

Adoption and governance belong inside the value system, not beside it. People will not reinvent their work while the metrics and rewards around them still enforce the old way, so leaders have to change those conditions rather than merely exhort. Governance, in turn, is better read as value infrastructure than as compliance: clear permissions, human-review thresholds and escalation routes are what let AI move into consequential work safely, and weak controls strand it in the low-value tasks where it can never earn much.

A short set of questions separates a real value case from a hopeful one. Are we solving a model problem or an operating-model problem? Which workflow constraint will this relieve, and who owns that process? What data must become reachable, and which decisions may the system change? How will released capacity be converted into a business result, and which metric will prove it? And what would cause us to stop? A programme that cannot answer these is measuring activity and hoping it becomes value.

None of this lessens the technology. Models will keep improving, and that progress matters. But the next wave of AI value will come less from models becoming more capable and more from organisations becoming more capable of using them. Capability sets what is possible. The operating model decides how much of that possibility the business ever sees.


Click Here to Schedule an Analyst Chat

Links:


Originally posted on 2025-01-07 in the IRPA AI Network — Announcements & Updates