Author: IRPA AI Senior Analyst, Kieran Gilmurray
This article explores a practical shift in enterprise AI. Agents are no longer being positioned just to answer questions. They are increasingly being positioned to carry out multi step work across tools, files, and systems. Once that happens, the management questions stop looking like prompt engineering questions and start looking like leadership questions: role clarity, permissions, oversight, escalation, and accountability.
After reading this article you will understand what an AI agent is in operational terms, why identity and authorisation are now the real scaling bottleneck, what new risks appear when agents can act across systems, and what a minimum viable governance package looks like for UK, EU, and US organisations in 2026.
What an AI agent is, and why that matters now
The simplest practical definition is this: a chatbot responds, a copilot assists, and an agent takes action. That action may still be narrow and supervised, but it is the change that matters. Once a system can search, call tools, write into applications, or complete a sequence of steps over time, it starts to behave less like a helper and more like a delegated operator.
That is why the “treat agents like team members” framing is useful. It does not mean pretending software is human. It means using familiar management disciplines. If an agent has no role description, no approval boundaries, no audit trail, and no escalation path, then the organisation is not scaling automation. It is scaling ambiguity.
What changed in March 2026
Microsoft described Copilot Cowork as long running, multi step work inside Microsoft 365, grounded in Microsoft 365 context and designed with auditable actions, sandboxing, and default identity, permissions, and compliance controls. It also positioned Agent 365 (a central control layer used to monitor, secure, and manage AI agents across an organisation) as a control plane for observing, securing, and governing agents at scale.
OpenAI’s March API changes point in a similar direction. These are updates to the developer platform that powers many enterprise AI tools, and they introduced capabilities such as tool search, computer use (allowing AI to operate software), and context compaction (allowing longer running tasks). Together, these make it more practical for AI systems to carry out multi-step workflows rather than just respond to prompts. Google’s late February update to its Agent Development Kit ecosystem makes the same strategic move explicit, describing a shift from chat to action and expanding integration paths, including MCP based connectivity.
That combination matters because it changes the enterprise question. Leaders are no longer asking only whether these systems are clever enough. They are asking how fast they can be deployed across functions. That is exactly the moment where operating model weakness becomes dangerous.
Why identity and permissions are now the scaling bottleneck
The most important governance shift is that agent scaling depends less on prompt quality and more on identity and authorisation. NIST (the US National Institute of Standards and Technology) and its applied research arm NCCoE (National Cybersecurity Center of Excellence) have been working on this problem. Their 2026 concept paper focuses on how organisations should identify, authenticate, authorise, and audit AI agents, and how to manage risks like misuse or hidden manipulation.
That is the right framing. A shared API key is not an operating model. If an agent can retrieve files, update records, send messages, or trigger transactions, then leaders need to know exactly which identity performed the action, on whose behalf, with which permissions, under which policy, and with what evidence trail.
This is why “agents as team members” should translate into role design. Every agent should have a defined objective, allowed actions, forbidden actions, data boundaries, approval rules, and a human owner. Without that, scale just means more invisible authority spread across the organisation.
The security problem gets worse when agents can read and act
The major security warning in this cycle is not abstract. Researchers have shown that AI systems can be influenced by hidden instructions embedded in web pages, emails, or documents. This is often referred to as prompt injection, but at a leadership level the important point is simple: content an agent reads can change what it does.
That changes how leaders should think about web access, email ingestion, ticket queues, and document processing. These are no longer just information inputs. They are potential control risks because they can influence system behaviour.
The safe assumption is that untrusted content is hostile by default. If an agent reads the web, inboxes, support tickets, or uploaded files, then leaders should assume those channels can carry hidden instructions and design controls accordingly. This means isolating instructions, restricting what actions agents can take, monitoring behaviour, and requiring approval for high impact decisions.
What “treat agents like team members” should mean in practice
Start with role clarity. A human employee normally has a job description, a line manager, a scope of authority, escalation routes, and performance expectations. Agents need the same logic. They should not be deployed as general purpose systems across the organisation. They should be deployed as bounded digital roles with explicit tasks and boundaries.
Then design supervision that scales. Human oversight does not mean manually checking every step forever. It means defining where approval is mandatory, where monitoring is enough, and where autonomy is acceptable because the action is low impact and reversible. The practical model is tiered: suggest only, act with approval, then act within narrow limits.
Finally, treat audit trails as part of the system. High value agent deployments should log intent, approvals, tool calls, data access, and outcomes in a way that can be reviewed and validated later.
A minimum viable governance package before scale
Leaders do not need a perfect framework before they start, but they do need a minimum viable package (the smallest set of controls required to operate safely in real conditions) that can function under real workloads.
First, require an owner for each agent. Not a vendor or a platform, but a named human accountable for outcomes. Second, require a role description for each agent: objective, tools, data sources, approval rules, and stop conditions. Third, require unique identity and revocation capability so the organisation can suspend access quickly when needed.
Fourth, define approval gates for high impact actions. Money, customer commitments, data deletion, access changes, and decisions affecting individuals should not be left to broad autonomous execution. Fifth, require auditable logs and monitoring for abnormal behaviour. Sixth, maintain a stop process: pause, revoke, terminate, and decommission.
This is also the point where procurement becomes governance. If a supplier cannot explain monitoring, audit support, model and tool versioning, identity integration, and termination processes, then the product is not ready for high consequence workflows.
UK, EU, and US differences leaders need to plan for
EU: The AI Act timetable is the main anchor. Enforcement and obligations increase through 2026, with broad application from August 2026. Leaders need documentation, oversight, and evidence in place before enforcement pressure arrives.
UK: Regulatory direction is still evolving. The ICO has signalled further guidance on automated decision making and profiling, and government analysis emphasises monitoring, oversight, and accountability as AI systems become more autonomous.
US: In federal contexts, the focus is on lifecycle governance, monitoring, and documentation. Public sector directives are pushing organisations to demonstrate control, evaluation, and accountability across the full system lifecycle.
Across all regions, the direction is consistent. Stronger governance and clearer accountability are becoming expected.
A practical rollout path
The right rollout path is bounded first, broad later. Start with a small set of use cases where value is clear and risk is manageable. Define the role, apply least privilege, and instrument the workflow before expanding autonomy. This keeps the first deployments narrow enough to govern properly and useful enough to learn from.
Then test agents in real workflows, not polished demos. Include misuse scenarios, edge cases, and failure modes, and evaluate behaviour under realistic conditions before expanding scope. Safe scaling depends on proving that governance travels with the system, not just that the system works in isolation.
Conclusion
AI agents are moving from conversation to action, and that shift changes the problem from technology to management. Once systems can act across workflows, value depends on how well organisations define roles, control permissions, monitor behaviour, and maintain accountability. The organisations that treat agents as managed digital roles with clear boundaries and governance will scale safely, while those that do not will find that the real constraint was never capability, but control.
About the Author: Kieran Gilmurray
Kieran is a globally recognized authority on AI, automation, and digital transformation, having authored multiple influential books and hundreds of articles that have earned him prestigious accolades, including being named a Top 50 Global Thought Leader and Influencer on Generative AI in 2024, a Best LinkedIn Influencer for AI and Marketing, Top 50 Global Thought Leaders and Influencers on Manufacturing 2024, Top 14 people to follow in data and one of the World’s Top 200 Business and Technology Innovators.
CLICK HERE TO SCHEDULE AN ANALYST CHAT
Links:
Originally posted on 2025-01-28 in the IRPA AI Network — Announcements & Updates