CJ Moses, Amazon’s chief information security officer, told the Fal.Con 2026 audience that Amazon’s MadPot honeypot network had captured an AI agent completing a full attack autonomously. “One of the honeypots actually captured an AI agent that was completing a full cyber attack, and it did it in 12 minutes and 42 seconds,” Moses said onstage with CrowdStrike President Michael Sentonas. “It was 94 events, zero syntax errors, the ability for it to respond in less than 500 milliseconds.”
The speed forcing the choice is real. Adam Meyers, CrowdStrike’s SVP of Counter Adversary Operations, told the Fal.Con 2026 audience that in the prior month CrowdStrike had seen almost as many agentic adversaries as it had in the six months before that. CrowdStrike reported an 89% year-over-year increase in attacks by adversaries it classified as AI-enabled, according to its 2026 Global Threat Report. “I will argue that the breakout time is over,” CrowdStrike President Michael Sentonas told the Fal.Con audience. “We now live in a world where a model finds a vulnerability and weaponizes it at the same time.”
Together, three recent launches and Microsoft’s established hybrid represent architecturally different answers to that acceleration. CrowdStrike launched purpose-built defender models and Google released a security-tuned Gemini variant days apart in September. Palo Alto announced native Cortex support for selected Claude and Gemini models in June. Microsoft has run a hybrid approach since 2023. Should the model defending your environment be purpose-built on attacker data, or should it be a frontier model wired into a security platform? The answer a CISO picks this quarter is a multi-year architecture commitment.
The four vendors publish four different kinds of evidence, and no common unit exists to compare them.
But the market may not be choosing the way any of these vendors expect. In VentureBeat's July Pulse Research survey of 116 enterprises, 92 of the 93 running or piloting agents named a primary security layer — and 85 of those named a control that shipped with their model provider or cloud. Across the full sample, CrowdStrike appears in 7% of security stacks and Palo Alto Networks in 6%. The purpose-built thesis is a supply-side argument about who can build defender models. The demand side says most buyers are not crossing that moat; they are defaulting to whatever their provider already ships.
Who owns the adaptation layer
The four bets fall along a single design question. CrowdStrike and Google train or adapt their own models on proprietary data, Palo Alto harnesses frontier models built by someone else, and Microsoft combines security-specific capabilities with frontier-model services. All four build on general-purpose foundation-model technology before adding security-specific adaptation, context or orchestration: CrowdStrike post-trains Nvidia Nemotron, and Google adapts its Gemini Flash foundation model for cyber tasks. The axis that matters is who owns the security-specific adaptation and who controls when the underlying model changes.
CrowdStrike launched SafeMind on September 1 at Fal.Con with Nvidia CEO Jensen Huang onstage. The system pairs two purpose-built models running on Nvidia Nemotron, post-trained with Falcon sensor telemetry, threat intelligence, and incident-response annotations drawn from 15 years of operations. Dr. Bartley Richardson, CrowdStrike’s chief AI and autonomous systems officer, said in his Fal.Con 2026 keynote that the training corpus includes 3.1 million working hours of expertise from what he called Falcon Complete MDR world-class detection engineers. CrowdStrike’s internal evaluations claim 29% higher detection, 6x faster remediation, and 99% lower cost against rival frontier and open-source models.
Google released Gemini 3.8 Flash Cyber on September 2, a security-tuned variant of its flagship Flash model built for autonomous vulnerability discovery and patching. Google reported 86.2% on CyberGym for vulnerability discovery. On CWE-Bench, an external patching benchmark administered by Collinear, it reported 47.2% pass@1. Google’s Chrome Security team reported that 3.8 Flash Cyber produced 2.6 times more correct patches to Chrome vulnerabilities than larger commercial models. Access is available through the Fairwind Program, which prioritizes trusted government authorities, critical infrastructure operators, and software maintainers.
Palo Alto Networks emphasized a platform-led, third-party-model strategy. In June, the company announced private-preview native Cortex support for Claude Sonnet 4.6, Claude Opus 4.8 and Gemini 3.5 Flash, building an integration layer rather than a proprietary foundation model. When Palo Alto validates and adds a new model to its approved Cortex set, customers can adopt it without retraining a foundation model, though they still need to validate workflows and guardrails.
A frontier-model harness creates a data-governance question that buyers need to resolve. The Palo Alto launch and release materials reviewed by VentureBeat describe native telemetry and governed agentic workflows, but they do not specify the data-retention, processing-region, inference-routing, or training-use terms that apply when Cortex invokes a third-party model — terms a buyer would need in writing before signing. Purpose-built model architectures sidestep this question by keeping inference inside the vendor's own pipeline.
Microsoft positions Security Copilot as a hybrid of security-specific capabilities and Azure OpenAI-based frontier-model services. In a randomized controlled trial, Microsoft reported that analysts using its phishing triage agent identified up to 6.5 times as many true positives per analyst minute and improved verdict F1 score by 77% against a control group. The study attributes 83% of that productivity gain to the agent reordering the queue and auto-resolving reports it judged benign, rather than to its verdicts — a workflow result more than a model one.

CrowdStrike says fewer than five companies can do this
In a press roundtable at Fal.Con with VentureBeat and other outlets, a journalist asked how many companies could build the full stack from model to harness to platform to ecosystem, suggesting the number was roughly five. George Kurtz, CrowdStrike’s CEO, went further. “Less than five,” he said.
“This isn’t a copilot that’s baked into someone else’s intelligence,” Kurtz said in his September 1 opening keynote. “It is a frontier-class model built and trained by CrowdStrike on our data in partnership with Nvidia.”
SafeMind is CrowdStrike’s purpose-built agentic AI system for cybersecurity, developed by the Cyber Superintelligence Lab. It pairs two specialized models in what CrowdStrike calls adversarial coevolution. Red Tempest, the offensive model, emulates AI adversaries to find attack paths across a digital twin of the customer’s environment. Blue Solano, the defensive model, writes detection rules for those attack paths.
Richardson put numbers on the economics in his Fal.Con session. Red Tempest achieved compromise of a test environment for $21 per task, compared with $96 for a closed frontier model and $62 for an open one in a harness. Blue Solano writes a detection rule for three cents, versus $10 with an off-the-shelf model. Compute runs on CoreWeave infrastructure. “We just commoditized defense,” Richardson said. CrowdStrike is the only vendor in this comparison to publish cost-per-task economics. Those figures have not been independently audited. No other vendor provided comparable data.
Kurtz drew a line between CrowdStrike and the frontier providers. “If you look at the frontier model providers, they’re like, hey, we solve a lot of problems, and then we’re going to go solve coding. Great, they’re going to solve other things. Our whole focus is going to be we need a model for defenders. Like, they don’t have one that works,” he told the roundtable.
AJ Shipley, CrowdStrike’s chief product officer, told VentureBeat in an exclusive interview at Fal.Con that the company is “not model agnostic” but supports “an open ecosystem of harnesses and models.” The SafeMind harness routes requests to CrowdStrike’s models or to Anthropic, OpenAI, or open-source alternatives depending on the task. Shipley said endpoint-based agent activity generated more telemetry than browser activity, the sensor reach CrowdStrike's purpose-built case rests on. He did not quantify the difference.
Not everyone in the security community accepts the threshold. Kayne McGladrey, an IEEE Senior Member and independent vCISO not affiliated with any of the four vendors in this story, argued that the barrier Kurtz describes is real but narrower than it sounds. "The barrier to entry is normalized, labeled data and funding allocation, not some magical capability gap," McGladrey told VentureBeat. "Companies like Fortinet, Cisco, Darktrace, SentinelOne, and others certainly have enough data and capital to train foundational defender models if they wanted to."
McGladrey offered a blunter test for CISOs weighing the decision. Most victims of AI-driven attacks, he said, "were losing to basic cybersecurity hygiene failures long before the swarms of agents arrived." His advice was to identify key business systems and the technical risks to those systems, and to close those gaps before buying the next generation of AI tooling.
In an exclusive interview with VentureBeat during Fal.Con week, Andrew Obadiaru, CISO at Cobalt and independent of every vendor in this story, urged security leaders to map what their agents can access and do before increasing their autonomy.
His guidance was specific. “Before even increasing an agent’s autonomy, it is important to map its identity, permissions, tools, and data, and what the downstream actions they have access to,” Obadiaru said. “If you don’t know what you have, you can’t protect against it.”
Obadiaru recommended adversarial testing and observability to examine what an agent can do across its trust and action chain. He cautioned against relying on a single tool to establish that assurance.
Five questions before the 2027 budget locks
Each of the four vendors measures something different and reports it differently. What a CISO needs before committing a 2027 SOC budget is a common framework — five questions that apply to all four platforms equally.
First, clarify what the vendor’s published metrics and data actually measure. A detection uplift inside one platform, a patching score on an external benchmark, and an analyst productivity trial are three different units of work. Treat them as three different answers, not three entries on the same scoreboard.
Second, identify what training data the model is going to see and how customers can verify the model is only using that data. A purpose-built model may incorporate proprietary attacker telemetry and annotations that a frontier model does not receive as part of its base training. Frontier models may bring broader general-purpose reasoning and coding capabilities. Buyers need task-level evidence to know which advantage carries into their environment.
Third, what happens to detection rules and playbooks if you decide to switch in two years? A purpose-built model can give the vendor more control over the training and inference pipeline. A frontier harness can shorten the path to approved model upgrades, though the platform vendor still decides which models are validated, supported, and available in each region. Switching costs include contract termination fees and the engineering effort to move detection logic, prompt chains, case context, agent permissions, telemetry mappings, and regression tests to another platform.
Fourth, how quickly do existing customers get new model versions, and how substantial are those revisions? How transparent are the vendors on their roadmaps and point releases, and are they asking for feedback in advisory councils? Get the update commitment and the notice period for breaking changes in writing.
Fifth, does their architecture allow for exporting a working production workflow and reproducing it elsewhere without rebuilding from scratch? How transportable are the workflows, projects, skills, and advanced logic created for a specific strategy? In previous architectures, this is where the lock-in led to being tied to a vendor long after their system couldn’t keep up with a growing business. CISOs should ask for references from customers who have successfully exported production workflows and rebuilt them elsewhere.
Across all five questions, the model is only one layer of lock-in. The harder asset to move is the operational security system built around it. Detection rules, prompt libraries, investigation context, identity boundaries, enrichment pipelines and automation playbooks accumulate faster than most teams realize.
Production sorts this out
CrowdStrike built SafeMind on the thesis that sensor reach, annotation depth, and incident-response history separate the companies that can build purpose-built defender models from those that cannot. Google’s security research and model-development resources position it to pursue the same approach. Microsoft took a different route, combining security-specific capabilities with frontier-model services and backing them with a published controlled trial. Palo Alto bet that foundation models will keep improving fast enough that owning the adaptation layer and the governance controls matters more than owning the model.
The Pulse data suggests the market may settle this before the vendors do. Most enterprises already run provider-native controls as their primary layer. The specialists have to displace that default before the architecture argument matters.
Which bet pays will be settled in production over the next two years, and the 2027 SOC budget is where most CISOs have to pick a side.
