If you sit in enterprise boardrooms today, the pressure weighing on CIOs comes down to two things: Velocity and ROI. Boards expect a mature agentic AI strategy and a definitive timeline for deploying autonomous systems that deliver measurable business impact.
In the rush to production, a critical architectural blind spot has emerged. The prevailing market narrative rests on a simple formula. Agents, plus tools, plus data are enough. While no vendor explicitly states it, nearly every product pitch assumes it. The idea is that you drop your corporate data into a cloud data lake, wire up an off-the-shelf large language model (LLM), hand it a toolbox of enterprise actions, and agility follows.
It won’t. It does the opposite. Having spent my career building developer platforms and cloud infrastructure for some of the world's largest workloads, I have learned that when enterprise software touches employees, customers, and finance, the physics change. There is zero tolerance for hallucination or drift when human livelihoods and corporate balance sheets are on the line.
The core problem of the lawless agent
To understand this friction, look closely at an agent's anatomy. At its foundation, an LLM is probabilistic, generating tokens based on mathematical likelihoods rather than hard-coded rules. The tech industry’s primary answer to introducing rules is tool-calling, which gives the model a set of APIs to act on in the real world.
While tool-calling makes a single action deterministic – where the tool reliably executes a defined instruction the same way each time – the overarching decision layer remains probabilistic. The agent must still decide which tool to call, when, and with what arguments.
Because a generic tool executes whatever instructions it is handed, tool-calling doesn’t solve the governance problem. It shifts the problem to the tool boundary, which is left exposed and ungoverned, a risk category that enterprise teams are increasingly expected to manage through formal AI risk management frameworks and open, secure AI infrastructure.
Wrap that probabilistic model in an agentic framework, assign it an objective, and you create a highly capable, autonomous, goal-seeking entity, with one flaw. There’s no native concept of corporate governance.
Consider a manager asking an off-the-shelf agent to onboard an international contractor. Operating in a structural vacuum, the agent will optimize entirely for task completion. Every individual tool call might succeed, but along the way, the agent can bypass regional labor laws, tax withholdings, and worker classification checks.
It does this not out of malice, but because nothing in its architecture understands that these rules exist. A lawless agent completes tasks through whichever systems aren’t paying close attention to the rules. On the surface, it looks like an agility win. In reality, it is a ticking regulatory time bomb.
The engineering instinct and the shadow ERP trap
When IT leadership catches this lawless behavior, the default engineering instinct is to patch the vulnerability from the outside. Teams resort to building external LLM gateways, layering on policy middleware, or stuffing massive corporate policies into an already overcrowded context window.
None of these fixes hold. A prompt is a request, not a guarantee. Fast-forward 18 months, and your engineering teams will have spent millions manually reconstructing organizational structures, routing logic, and security permissions inside a makeshift, isolated AI stack.
This is the “shadow ERP” trap, where enterprise agility goes to die. You end up maintaining a fragile, buggy copy of the same business logic you already owned inside your system of record. It’s no shortcut when it leads to endless maintenance costs and fragmented governance.
The solution: Guardrails at the inference layer
To make an agent lawful by design, enterprise safety must become core infrastructure, wired directly into the inference engine and the execution runtime. Fully controlling what an agent can see, do, and touch requires running AI inference on top of the platform that already provides enterprise governance:
A context engine: Assembles precisely what the agent’s identity is entitled to see – including active records, current policies, and organizational context. This filters the agent’s entire reality, ensuring its context window and exposed database rows are pre-governed.
A policy engine: Validates every agentic action against the exact business logic that governs a human counterpart. If a financial transaction requires four levels of executive approval, the platform refuses every path around them.
An execution layer: Ensures every external API call carries the explicit authority and scope of the specific human the agent serves, rather than a standing “god-mode” credential. The agent cannot exceed its authority because the capability was never issued.
The truly difficult part of enterprise AI isn’t reasoning; it’s compliance. A single payroll run sits on top of shifting withholding and wage rules across thousands of global jurisdictions. No standalone model can recreate that machinery. Your enterprise already knows this, which is why you invested in a system of record rather than writing a custom payroll engine. Routing agents through external middleware quietly reverses that foundational decision. If you want a lawful agent, you must run it where the laws already live.
The proximity principle
Every critical guardrail depends entirely on live enterprise state – the organizational graph, active entitlements, and business processes as configured. If you attempt to rebuild these elements from the outside, the physics of software will catch up with you. What the agent sees degrades into a lagged snapshot, and what it touches relies on risky, static API keys. Distance downgrades guardrail integrity.
Distance also costs speed. Agentic loops are inherently chatty, requiring dozens of tool calls to complete a single task. Every time a call crosses an external trust boundary, it must re-prove who is asking. Inside the native platform's perimeter, that security context is ambient. The architectural choice that makes agents lawful is the same choice that makes them fast.
Call it the “proximity principle.” AI inference belongs directly on the platform that governs it, because both guardrails and operational speed decay with every hop in between.
When the next agentic pitch reaches your desk, look past the demo and ask one question: Where do the guardrails run? If the answer is an external layer your own team must build and maintain far from your system of record, you are building a shadow ERP. To achieve true velocity, anchor your agents directly to the system that already holds your enterprise laws.
Gabe Monroy is CTO at Workday.
Welcome to the VentureBeat community!
Our guest posting program is where technical experts share insights and provide neutral, non-vested deep dives on AI, data infrastructure, cybersecurity and other cutting-edge technologies shaping the future of enterprise.
Read more from our guest post program — and check out our guidelines if you’re interested in contributing an article of your own!
