Author: IRPA AI Senior Advisor, Kieran Gilmurray

Ask most executives how their AI programme is going and you will get a number about people. Seats deployed, weekly active users, hours saved per employee. The numbers are usually good, and they are usually true. Microsoft’s early Copilot research found users completing tasks faster and feeling more productive, and there is no serious reason to doubt it. The awkwardness comes when the same executive is asked what any of it did to the P&L.

McKinsey’s 2025 State of AI survey tested twenty five attributes against enterprise EBIT impact from generative AI, and workflow redesign had the single biggest effect of any of them. Only 21% of organisations using gen AI say they have fundamentally redesigned even some workflows, and more than 80% report no tangible enterprise level impact. This article is about the space between those two facts: why individual gains keep failing to become organisational value, where exactly the value leaks out, and what leaders have to redesign before the next wave of spending makes the problem more expensive.

The Productivity That Doesn’t Arrive

Gartner’s supply chain research contains the most uncomfortable finding in enterprise AI, and it deserves to be read slowly. Desk based workers using generative AI saved 4.11 hours per week. At team level, that figure fell to 1.5 hours per team member. And those team level savings showed no correlation with improved output or quality.

Nothing was faked. The individual hours were genuinely saved. They simply did not survive contact with the system the individual works inside. Some were absorbed by waiting for someone else. Some went into rework caused by a handoff that AI made faster but no more accurate. Some were reinvested in more of the same work rather than different work, because nobody had decided what the freed capacity was for. The gain was real at the desk and gone by the department.

Gartner found the same shape in finance, where only 34% of teams using generative AI reported high productivity gains, marginally below teams using traditional AI, which is why the firm told CFOs to reset expectations and redesign workflows and structures instead. EY’s Work Reimagined survey puts the endpoint plainly: 88% of employees now use AI at work, and 28% of organisations are positioned to turn that into high value outcomes. Access is close to universal. Conversion is not.

Why the Task Is the Wrong Unit

The copilot model contains an assumption that nobody stated out loud because it seemed too obvious to state: that if you make every task faster, the process gets faster. This holds only when the process is a queue of tasks performed by one person. Almost no valuable enterprise process looks like that.

Real processes are sequences of handoffs, approvals, waits, checks, and rework loops, and the time cost sits overwhelmingly in the joins rather than the tasks. Speeding up a task inside a process governed by a weekly approval meeting produces exactly nothing. The work simply arrives at the bottleneck earlier and waits longer. This is why Bain’s formulation lands so hard: automating mediocre processes accelerates mediocre outcomes. The acceleration is real. It just points at the wrong thing.

The consequence is that the unit of change has to move from the task to the end to end flow, and this is not a semantic upgrade. It changes who owns the work, what gets measured, and where the budget sits. A task belongs to a person. A workflow belongs to nobody in most organisations, which is precisely why it never gets redesigned.

The Leak Map

If value is generated at the task and lost before it reaches the enterprise, the practical question is where. The Leak Map names five points where it goes, in the order it usually goes there. It is a diagnostic rather than a maturity model: most organisations are leaking at all five, and fixing the fifth without the first is theatre.

The handoff leak. Time saved upstream is surrendered at the transfer to the next team, function, or system. The faster the upstream step, the more conspicuous the wait.

The decision leak. The work arrives at a human decision that has no defined threshold, no delegated authority, and no service level. AI compressed the analysis from three days to three minutes, and the decision still takes a fortnight.

The rework leak. Output produced faster but validated no better returns for correction later, usually at a more senior and more expensive point in the process. Volume up, quality flat, cost migrating upwards.

The absorption leak. Freed capacity is reinvested in the same work rather than redeployed to higher value work, because nobody made an explicit choice about it. This leak is invisible in every dashboard, which is why it is the largest.

The measurement leak. The organisation counts usage rather than outcome, so the other four leaks never appear in any report and the programme is declared a success while the EBIT line does not move.

The map’s usefulness is that it converts a vague complaint about ROI into a locatable fault. Ask which of the five is costing you most in your highest value process, and the conversation stops being about AI and starts being about how the firm actually works.

What Redesign Actually Produces

The counterfactual matters here, because it would be easy to conclude that AI simply underdelivers. The organisations that redesign end to end are getting results of a different order, not a better version of the same order.

BCG’s work on agentic operations draws the line sharply. First wave deployments that layered AI onto existing work generated something in the range of 10% to 20% productivity improvement. Early agentic redesigns, where the process itself was rebuilt, are showing threefold productivity gains, cycle time reductions around 80%, and long term cost reduction of 60% or more. BCG describes a European bank that redesigned retail lending holistically rather than augmenting it, reaching more than 90% end to end automation on consumer loans and more than 70% on mortgages. IBM’s internal Client Zero programme reports over 100 AI enabled workflows and 4.5 billion dollars in productivity gains, and the operative word in that sentence is workflows.

Salesforce’s account of its own service function is the most instructive because it is the least flattering. The agent handled over 1.5 million support requests, the majority without human involvement, but getting there required revising goals, cleaning and consolidating contradictory knowledge, integrating across CRM, Slack, web and email, and redesigning human roles around complex cases. None of that is AI work. All of it is the work.

The gap between the 10% to 20% band and the threefold band is not explained by better models. Both groups have access to the same models. It is explained entirely by whether the organisation was willing to change the shape of the work.

Why Agents Make This Urgent

The copilot era was forgiving. A copilot suggests, a human decides, and the process debt underneath stays hidden because a person is standing in every gap absorbing the inconsistency. An agent executes, routes, and coordinates across systems, which means it runs directly into the handoffs, the missing decision rights, and the contradictory data, and it does so at speed and at scale.

The market is already moving on this basis. Gartner expects that by 2028 more than half of enterprises will stop paying for assistive intelligence alone and will favour platforms that commit to workflow results. Microsoft reports 46% of leaders already using agents to fully automate workflows or processes, and 81% expecting agents to be moderately or extensively integrated into their AI strategy within 12 to 18 months. KPMG finds 73% of organisations using agents to automate workflows spanning multiple functions, with most requiring human validation of agent outputs, which is worth naming for what it is: supervised autonomy, not delegation.

The scepticism is warranted, though, and should be held alongside the ambition rather than instead of it. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 on grounds of unclear value, cost, or inadequate controls. ServiceNow finds 59% of organisations past the agentic pilot stage but only 9% making meaningful progress on autonomous multistep workflows, with AI enabled workflows scoring 40 out of 100, the lowest pillar in its maturity index. These are not signs that the direction is wrong. They are signs that orchestration is difficult and that firms are buying it faster than they are building the conditions for it to work.

The Human Half of the Redesign

There is a version of this argument that reads as though the goal is to remove people from the flow, and the evidence does not support it. PwC’s 2026 barometer finds productivity growth 40% higher at the companies most exposed to AI, and finds that new tasks added to AI exposed roles are 2.5 times more likely to depend on judgement, empathy, creativity, and leadership. The redesign moves people, it does not delete them, and it moves them towards the parts of the process that were always the hard parts.

Microsoft’s framing of a shift from org chart to Work Chart, where teams form around outcomes rather than functions, is the organisational expression of the same idea. If the unit of value is the workflow, then the unit of team design has to be the workflow too, which is uncomfortable for anyone whose authority derives from a function. EY’s finding that employees embrace AI more readily when they understand how it is used and retain override control tells you the sequence: intent set by humans, execution bounded by policy, exceptions returned to judgement, accountability unmoved.

What This Means for Leaders

The programme review most executive teams should run is not about AI at all. Take the two or three workflows that carry the most value in the business, walk them end to end, and find the leaks. Where does time surrendered at handoffs? Which decisions have no threshold and no owner? Where is faster output creating slower rework? What happened to the capacity you released last year, and can anyone say? If the answer to the last one is a shrug, the AI investment has been funding activity rather than performance, and adding agents will fund more of it.

Then decide three things before the next tranche of spending. Which one or two end to end workflows get redesigned first, chosen for value rather than ease. What humans still decide, written down as policy rather than assumed as culture. And what you will measure: cycle time, exception rate, quality, conversion, margin, not seats and prompts. Bain’s advice to pay down workflow debt and name a governance owner before scaling is the least glamorous item on any board agenda and probably the highest returning.

Technology creates possibility. Management creates value. Every competitor will have the same copilots by the end of the year. The advantage belongs to whoever redesigns the work they sit inside.


Links:


Originally posted on 2025-01-07 in the IRPA AI Network — Announcements & Updates