After studying more than 100 internal AI transformation efforts, Microsoft says the winning enterprise architecture will put proprietary evals, context and learning loops above interchangeable foundation models — while redesigning entire workflows before handing them to agents.

For enterprises racing to deploy AI agents, Microsoft has a counterintuitive message: Don’t start with the agents.

The company’s new 44-page “Becoming a Frontier Firm: Our Frontier Playbook,” being released Thursday, argues that many organizations are making the same mistake Microsoft says it made early in its own AI transformation — treating AI like another technology rollout, measuring licenses and usage, and dropping new tools into workflows that were designed for humans and legacy software.

Instead, Microsoft says enterprises should first dismantle and redesign the underlying process, establish a shared data foundation, decide explicitly which decisions humans and agents should each control, and only then introduce agents.

And perhaps more consequentially for enterprise developers, Microsoft argues that companies should avoid making any particular foundation model the center of that architecture. Instead, their durable competitive advantage should reside in private evaluations, proprietary context, workflow orchestration, feedback loops and institutional knowledge that can survive the replacement of the underlying model.

That amounts to a strikingly model-agnostic vision from one of the world’s largest AI platform vendors.

Microsoft developed the playbook after reviewing more than 100 internal AI transformation case studies across its corporate functions, commercial organization and engineering teams, according to the document. Its central message is that simply giving employees AI tools is not transformation.

“A tool licensed and rolled out to 100,000 employees does not change how the work gets done,” the playbook states.

Kathleen Hogan, Microsoft’s chief strategy and transformation officer, makes the point even more starkly in an accompanying blog post: Microsoft initially approached AI like a traditional software rollout — deploy the technology, train employees and drive adoption — only to find that access and usage did not necessarily produce business impact.

For enterprise technology leaders, the most useful parts of Microsoft’s playbook are therefore less about Copilot itself than the operating architecture Microsoft says it is building around AI.

Microsoft’s rule: ‘Lean before agents’

Microsoft divides enterprise AI transformation into three approaches:

  1. Persona Acceleration, which gives particular roles AI tools tailored to their work

  2. AI-Powered Process Redesign, which reconstructs existing workflows around AI

  3. AI-First Possibility, in which teams start essentially from a blank sheet of paper and assume AI will be embedded from the beginning.

Its guidance for the second category is particularly applicable to organizations moving from copilots to agents.

Microsoft recommends mapping a workflow end to end, removing unnecessary approvals and handoffs, creating a shared data foundation, establishing cross-functional ownership, building reusable orchestration and observability infrastructure, and deliberately deciding which parts of the process should remain human-led.

The company calls the approach “lean before agents.” Microsoft says its own cloud supply-chain organization followed that pattern before deploying 111 purpose-built agents across planning, sourcing, fulfillment and logistics.

Rather than placing agents on top of the existing process, a cross-functional team first mapped and simplified workflows and created what Microsoft describes as a single source of truth from which the agents could reason.

The resulting agents can investigate changes in demand, model capacity and compare transportation alternatives across air, land and sea based on factors including cost, timing and carbon impact.

Within specified permission and approval thresholds, some have advanced beyond providing information and can help planners update or cancel purchase orders.

Microsoft says selected supply-chain workflows subsequently reduced average cycle time by as much as 75%. In five monthly planning cycles measured between April and August 2026, average cycle time declined from roughly 10 business days to less than 2.5.

For more than 20 demand-plan investigations conducted each month, Microsoft says producing a human-validated explanation for a change previously took five to seven days; it now takes less than several hours, with some investigations completed in under 20 minutes.

Those figures come from Microsoft’s own analysis of a 150-plus-person cross-functional effort conducted between September 2025 and August 2026, and the company cautions that the results apply to the particular workflows and measurement periods rather than constituting a general enterprise benchmark.

But the architectural lesson is more interesting than the percentage improvement: Microsoft is explicitly warning companies that adding agents to poorly designed processes may simply automate the dysfunction already present.

The playbook says the underlying shared data, orchestration, telemetry and governance layers can matter more than the visible agent itself.

From copilots to ‘human-led, agent-operated’ companies

Microsoft also proposes a three-level model for AI autonomy inside companies.

  • Level 1: employees work with AI assistants.

  • Level 2: agents become members of human-agent teams and take responsibility for particular tasks under human direction.

  • Level 3: Microsoft envisions what it calls “human-led, agent-operated” workflows, in which humans set direction while agents execute entire business processes and check in when necessary.

Moving up those levels requires more than better models. Microsoft says organizations need progressively greater data and infrastructure readiness, tool access, clearly defined risk boundaries and willingness to change human roles.

Software engineering could span all three simultaneously. A developer might use an assistant for one part of the job while delegating another workflow to an autonomous coding agent.

Microsoft offers one particularly aggressive example from its own product development organization. A nine-person team building what the document calls Copilot Cowork was placed in a sandbox and told to approach product development as AI-first rather than adding agents to an established engineering process.

The team adopted a spec-driven workflow built around three elements: specifications describing intent, evals defining what good output looks like and context supplied to the agents. Microsoft says team members evolved into what it calls “meta-engineers,” “meta-designers” and “meta-PMs,” working with agents from a shared context and common set of evaluations and agent definitions.

The playbook reports 18,600 commits, at a rate of 123 per day, and 9.3 million lines of code, with the nine-person team delivering an initial product release in 35 days.

The sheer code and commit counts should not be confused with product quality or developer productivity on their own. Microsoft itself notes separately that the 35-day result comes from one dedicated project and should not be interpreted as a companywide development benchmark.

What is notable is the operating model: rather than optimize individual developers to produce code faster, Microsoft restructured the team around agents, specifications, evaluations and shared context.

Microsoft thinks your AI moat should sit above the model

The playbook’s most consequential argument may concern something many enterprises are currently struggling to define: Where does proprietary advantage live when everyone can access increasingly capable foundation models?

Microsoft’s answer is that it should not live primarily in the model.

Instead, it encourages companies to codify their own definition of good performance through private, custom evaluations (evals) of what a "good" outcome looks like for any given business task assigned to an agent.

Microsoft breaks this proprietary intelligence into four categories: a company’s point of view about its market; proprietary data, workflows and institutional knowledge; its particular sense of quality or “taste”; and its risk boundaries governing where agents can and cannot act autonomously.

Those standards then become evaluations against which AI systems can be continuously measured.

However, in VentureBeat Intelligence's Agent Reliability and Evals Q2 Pulse Survey (July 2026 wave, 108 enterprises), just 13% said they fully trust automated evaluation today. Among the 53 enterprises that had already shipped an agent that passed internal evals and then failed in front of a customer, only 4% said they trust it — against 24% of the 41 enterprises that hadn't yet been burned

Microsoft proposes wrapping those evaluations in what it calls a “hill-climbing machine” — essentially a continuous learning architecture in which enterprise feedback, scoring and tuning improve AI systems against the company’s own standards over time.

The diagram Microsoft provides is revealing. Security and governance sit at the top. Underneath is a reinforcement-learning environment containing evals and rubrics; then agent runtimes, managed hosting and tools; then a context and “harness” layer encompassing multi-model access, enterprise knowledge, memory, skills and MCP; and finally the foundation models themselves.

The models are explicitly labeled “interchangeable.” That architectural separation is intentional. Microsoft argues that companies should own and protect the evaluation, context and control layers that make their businesses distinctive while remaining able to replace external foundation models as model quality, economics or requirements change.

In other words, Microsoft is telling enterprises not to confuse renting intelligence from a model provider with owning an AI strategy.

It also recommends keeping prompts, retrieval systems, evaluations, agent decisions and workflow intelligence within the enterprise boundary wherever appropriate, with architectural controls for data residency, tenant isolation, model independence and intellectual-property protection.

The playbook characterizes retrofitting residency and tenant isolation after deployment as vastly more difficult than designing for them upfront — another reason Microsoft argues that governance cannot simply be bolted onto agents after they have begun taking actions.

Measure the business, not the token count

Microsoft applies the same skepticism to AI measurement.

Its proposed framework separates indicators into four layers: inputs, such as adoption and readiness; throughput, showing whether AI is actually penetrating workflows; outputs, such as productivity or performance improvements; and finally business outcomes, such as revenue or customer success.

Microsoft suggests that these layers operate on different timelines: input changes can appear within two to four weeks, workflow changes in roughly one to three months, performance effects in three to six months and high-level business outcomes after six months or more.

That matters because many enterprise AI deployments are still judged primarily by seat adoption, prompt volume or developer usage.

Microsoft’s own sales experiment illustrates both the opportunity and the need for caution.

In a pilot involving 687 Microsoft 365 Copilot sellers during the first half of 2024, Microsoft says adoption of priority AI use cases tripled, revenue per account manager increased 9.4%, and deal close rates were 20% higher relative to sellers with low Copilot usage.

The comparison was based on internal observational data rather than a randomized experiment, so those figures should not by themselves be interpreted as proof that Copilot caused all of the improvement.

But Microsoft says its broader lesson was that pushing employees to “use AI more” was less effective than identifying concrete moments in their jobs where specialized agents could improve an outcome.

The foundation model may become the least durable part of the enterprise AI stack

Microsoft packages the playbook as a roadmap toward what it calls the “Frontier Firm” — a company that remains human-led while progressively shifting more operational work to AI.

It is also, of course, a Microsoft thought-leadership document produced by a company selling much of the infrastructure and software enterprises might use to implement that transformation. The internal performance results have not been independently validated, and Microsoft repeatedly notes that its experiences should not be assumed to generalize automatically to other organizations.

Still, there is an important architectural argument underneath the branding.

The first wave of generative AI encouraged companies to ask which model was smartest. The second encouraged them to put copilots into existing applications. Microsoft’s new playbook suggests the next phase will instead revolve around rebuilding the enterprise itself around agents — while deliberately ensuring that the model underneath those agents remains replaceable.

For enterprise developers and technology leaders, that shifts the strategic question.

The scarce asset may not be access to an increasingly powerful model. It may be the private evaluations that define what good means for a particular business, the proprietary context agents can reason over, the orchestration and security systems governing what they can do, and the feedback loops through which those systems keep improving.

If Microsoft is right, enterprises should be spending at least as much time building those layers as they spend choosing the next frontier model.