Chris Goettl spent months training a Claude skill on the same patch data he had processed by hand for a decade. It kept inventing details.

“You see something blatantly wrong, and at some point, you just get fed up, and you say, did you just make that up,” Goettl told VentureBeat. “And it will literally tell you that it has no actual foundation or referenceable material. That was all me.”

Goettl is Ivanti’s VP of product management for endpoint security. The skill he built applies Ivanti’s own threat-risk prioritization to raw patch data, runs on Anthropic’s platform, draws only on published vendor advisories and Ivanti’s own spreadsheets, and does not touch customer data. Every output is a draft until a human approves it for release. The prioritized Patch Tuesday briefing reaches 500 to 700 webinar attendees each month.

AI is inventing data in the pipelines that produce the security guidance enterprise customers act on. The largest security vendors now use AI to rank Patch Tuesday CVEs, and their customers do not know what produced that ranking. Ivanti is the only vendor VentureBeat found that has described fabrications from its own AI skill and the review step built to catch them.

Seven counts, 205 CVEs apart

On September 8, Microsoft shipped the largest Patch Tuesday in its history. Seven trackers published totals for the same Tuesday. Tenable counted 964 CVEs. Senserva, which counts CVEs against KB articles rather than advisories, reached 1,169. That is a 205-CVE gap for the same set of patches. Two of those vulnerabilities were zero-days already exploited in the wild before the fix shipped. The biggest Patch Tuesday before 2026, by Ivanti’s June post, was 175 CVEs in October 2025. September’s gap alone exceeds the old record.

VB seven counts chart

Microsoft stopped listing its CVEs. Every Patch Tuesday count is now a parse. CVEs resolved in Microsoft's September 8, 2026 Patch Tuesday, as published by each tracker. Sources: Tenable, BleepingComputer, Zero Day Initiative, Ivanti, Rapid7, SecurityWeek and Senserva, September 8 and 9, 2026. Chart: VentureBeat

Not every gap requires AI to explain. Scope definitions and counting methods produce a wide spread on their own. What AI changes is the layer between the raw data and the risk tier a customer acts on. Microsoft stopped presenting a single monthly CVE list in its Security Update Guide starting with July’s Patch Tuesday. Rapid7’s Adam Barnett documented the change on July 14. Zero Day Initiative’s Dustin Childs, whose independent ledger goes back two decades, opened his July review by declaring “the bug apocalypse has fully descended upon us.” Every vendor now builds its own list.

Inside the skill that produced the 973

Goettl walked VentureBeat through the skill in a recent interview.

He trained the system vendor by vendor. Microsoft first, because its downloadable spreadsheets made that training the cleanest. Adobe was harder. Adobe does not provide a clean spreadsheet, so Goettl fed the skill individual page URLs and trained it to read Adobe’s site structure. He configured the skill to flag deviations from each vendor’s release pattern, so known exploits, public disclosures, and abnormal CVE volumes all trigger the risk-tier escalation his customers depend on.

Adobe’s release cadence came back wrong repeatedly. Office editions needed clarifying because Microsoft breaks Office into multiple edition families and the skill conflated them. The pattern persists even in production. The fabrication catches are why the review step exists. Goettl built the human gate around the failure modes he found during training.

Goettl and Todd Schell, Ivanti’s senior product manager for patch, had spent about 48 hours around every Patch Tuesday researching vendors and assembling the prioritized briefing their customers receive. That effort consumed at least 5% of their combined working time each month.

Schell replayed the skill on two earlier months of hand-built data. Ivanti says its reviewers found 98% alignment, with a few edge cases to fix. VentureBeat did not independently audit the data set, scoring method, or error categories. The August run was the first production month. Goettl was on vacation. The spreadsheet that used to take about four hours each for two people ran in under 30 minutes. September 8, at 973 CVEs by Ivanti’s count, was the skill’s second production month.

pipeline graphic

From 48 hours of manual work to 30 minutes with a reviewer. Ivanti's skill-assisted Patch Tuesday pipeline, August and September 2026. Sources: Chris Goettl, Ivanti VP of Product Management, Endpoint Security, exclusive interview August 25, 2026. Chart: VentureBeat

The rest of the advisory chain

CrowdStrike and Palo Alto Networks both keep humans in the chain. At Fal.Con 2026, CrowdStrike described a “human on the loop” design where an analyst works the same detection in parallel with the agent, comparing verdicts on the same data at the same time. Palo Alto shipped Cortex XSIAM AgentiX in February 2026 with prebuilt agents gated by human-in-the-loop approval for high-impact actions. All three vendors converged on that design independently. Neither CrowdStrike nor Palo Alto has publicly disclosed a fabrication catch.

An independent read

Kayne McGladrey, senior IEEE member, independent vCISO, and author of “Cyber Risk is a Myth,” has no ties to Ivanti and answered in writing about the practice, not any one vendor.

“The vendor owes their customers one page, in writing, before a single priority task lands in anyone’s queue,” McGladrey wrote in an email to VentureBeat. “That page should cover where the model sits in the pipeline, whether deterministic code parses the CVEs or the LLM does the counting, who reviewed the output, and what fraction of a 400-row list got verified against the source rather than skimmed.”

Asked whether a replay plus a human read before release is sufficient, his answer was no.

“A replay checks whether this month looks like last month. It tells you nothing about correctness. A human reading it might not do much better, because the failure that matters is the CVE that never appears, and nobody skimming a machine-generated list will notice something’s missing.”

Schell compared the skill's output against months he and Goettl had built by hand, not against last month's shape. That comparison does not address McGladrey's second point. The risk tiers were not independently checked, and a row that never appears is still invisible to a reviewer reading the rows that did.

What ships and what holds

VentureBeat Pulse’s July 2026 Agent Reliability and Evals tracker surveyed 108 enterprises across its research panel. 49% had shipped an agent that passed internal evals and then failed in front of a customer. 37% let an autonomous agent deploy changes with no human in the loop. 13% fully trust automated evaluation. Among the 53 enterprises burned by an eval failure, just 4% still trust it.

trust collapse

Trust in automated evaluation collapses after failure. VentureBeat Pulse Agent Reliability and Evals tracker, July 2026, 108 enterprises surveyed. Chart: VentureBeat

Goettl’s team kept the human reviewer. CrowdStrike kept the parallel analyst. Palo Alto kept the approval gate.

The pressure driving adoption is real. CISA’s June 2026 Binding Operational Directive 26-04 set a three-day window for the highest-risk vulnerabilities at federal civilian agencies. CrowdStrike’s 2026 Threat Hunting Report found 88% of observed exploitation with a public proof of concept began inside 48 hours.

A disclosure no one reported

Ivanti published a Patch Tuesday post on June 9 captioned “Graph generated using Claude (Anthropic) on June 9, 2026, based on author-designed prompts and dataset by Chris Goettl.” No regulation or industry standard required the disclosure, and that caption has sat on Ivanti's site for three months.

Computer Weekly, Cyber Magazine, Help Net Security, Krebs on Security, and Rapid7 have all run Goettl’s numbers this year. None reported that the guidance is partly skill-produced, what the skill got wrong during training, or the review chain between it and release.

Three questions for any vendor shipping AI-assisted guidance

Goettl’s reasoning is that the judgment is the product. “What really matters is my subject matter expertise,” he told VentureBeat. “Knowing what’s important, what’s not, having that informed ability to make judgment calls, to give recommendations, to guide the tools to help make this happen.”

McGladrey’s one-page standard gives a patch team the language to hold any vendor accountable. Any SOC leader receiving Patch Tuesday guidance can ask three questions today and get an answer before October 13:

  • What in your Patch Tuesday guidance is AI-generated?

  • What was it validated against, and does that validation catch a missing row?

  • Who reviews it before release, and is that disclosed?

October 13 is the next Patch Tuesday. The parses begin again, and the total will depend on what each vendor chooses to count.