VentureBeat Intelligence surveyed respondents at organizations with 100 or more employees in its August VB Pulse survey. We asked about AI agent evaluation and production failures, whether agents may push changes to production with no human review, and which evaluation platforms enterprises run.
Among August respondents whose organization deploys autonomous AI agents, 56% already let an agent push some changes to production on automated evaluation results alone, or are building toward it. In July, 75% of respondents in that group gave one of those two answers. Over the same month, the share expecting to keep a human reviewing agents' production changes for the foreseeable future doubled, from 20% to 42%.
Among the same August respondents, 32% already allow it for specific low-risk agents or changes, and 24% are engineering their pipelines to allow it within 12 months.
Automated evaluations let problems reach customers. In this report, an evaluation-passing failure means an AI agent or LLM feature that passed internal evaluations and then caused a customer-facing failure. Leaving out the 5% of August respondents whose organization runs no pre-deployment evaluations, 61% reported at least one such failure in the past 12 months.
Yet among organizations that deploy autonomous agents, respondents reporting an evaluation-passing failure and those reporting none allow unreviewed pushes, or are building toward them, at rates too close to call.
Asked for the limitation that most reduces their trust in automated evaluation, only 9% of respondents chose the answer that they trust automated evaluation today. The rest named a limitation. Across all respondents, 2% of those reporting an evaluation-passing failure chose that answer, against 18% of those reporting none. Among organizations that deploy autonomous agents, the gap on that answer between respondents reporting a failure and those reporting none, 3% against 10%, is too close to call.
Few enterprises with AI agents build live monitoring around checking whether agent output is correct. Among respondents whose organization deploys autonomous agents, only 29% say their production monitoring is built around real-time automated quality checks. Trace logging or gateway tracking, neither worded as checking correctness, is the main approach for 53%.
A larger share of August respondents than July respondents run OpenAI's native evaluation tooling. Among respondents who answered the platforms question, 59% in August said their organization runs the OpenAI Developer Platform's native evaluations and traces, against 31% in July. In all, 62% of August respondents plan to adopt a new, additional or replacement evaluation platform within 12 months.
Finding 1. 56% of enterprises with AI agents let them push production changes on automated checks alone or are building toward it, down from 75% in July
The share expecting to keep a human reviewing agents' production changes for the foreseeable future doubled in one month, from 20% to 42%.
An automated evaluation is a set of tests that scores an agent's code or system change before release, and a pipeline can ship the change when the scores pass. Removing the human reviewer lets changes ship faster and frees engineers from routine approvals. Without a reviewer, a change that passes the tests but breaks something reaches production before anyone looks at it.
We asked respondents whether their organization would currently allow an autonomous AI agent to deploy a code or system change to production based on automated evaluation results alone, with no human-in-the-loop validation.
Finding 1 — 56% of enterprises with AI agents let them push production changes on automated checks alone or are building toward it, down from 75% in July
Base: 140 respondents. One answer each. Answers are worded as in the questionnaire, where "fully automated deployment" means an AI agent pushing a code or system change to production with no human review. The figures in the text leave out respondents who answered "not applicable" (Base: 118).
Finding 1 — Answer, among respondents whose organization deploys autonomous agents, July → August
Bases: 96 July and 118 August respondents whose organization deploys autonomous agents. One answer each.
Among August respondents whose organization deploys autonomous agents, 42% said their organization expects to keep a human reviewing agents' production changes for the foreseeable future. Among July respondents in the same group, 20% said so. The difference between the two months counts as a real change.
Another 56% of August respondents in the group already allow unreviewed pushes or are building toward them. Some already let an agent push a change with no human review for specific low-risk agents or changes. Others are actively engineering their pipelines to allow unreviewed pushes within 12 months. The remaining 3% did not know.
In July, 75% of respondents whose organization deploys autonomous agents already allowed unreviewed pushes or were building toward them. The difference between the two months also counts as a real change.
Final purchasing decision-makers whose organization deploys autonomous agents show the fall on their own: 88% already allowed unreviewed pushes or were building toward them in July, against 61% in August. Among other respondents in the group, the share went from 63% to 48%, a difference too close to call on its own. We did not test whether the change among final purchasing decision-makers differs from the change among other respondents.
Finding 1 — Allow it or building toward it, among respondents whose organization deploys autonomous agents, July → August
Each cell gives the share and count of that group in that month. Bases: final purchasing decision-makers, 48 July and 70 August; all other respondents, 48 July and 48 August.
Finding 2. 61% of enterprises that run evaluations shipped an agent or LLM feature that passed them and then caused a customer-facing failure
Most organizations that run pre-deployment evaluations shipped an agent or LLM feature that passed them and then caused a customer-facing failure, and 22% had more than one such failure.
Pre-deployment evaluations test an agent or LLM feature against a fixed set of cases before release. A test set covers only the cases someone thought to write, so a feature can pass every test and still fail on real customer requests. A failure in production costs wrong answers that customers act on, broken workflows and the engineering time to find and fix the problem.
We asked respondents whether, in the past 12 months, their organization had deployed an AI agent or LLM feature to production that passed internal evaluations but then caused a customer-facing failure. Examples given were an incorrect output, a broken workflow or a quality incident.
Finding 2 — 61% of enterprises that run evaluations shipped an agent or LLM feature that passed them and then caused a customer-facing failure
Base: 140 respondents. One answer each. The text leaves out respondents whose organization runs no pre-deployment evaluations (Base: 133 in August and 101 in July).
The 5% of respondents who said their organization runs no pre-deployment evaluations are left out of the figures below. Among the other respondents, 61% said their organization shipped an AI agent or LLM feature that passed its evaluations and then caused a customer-facing failure. Among the same respondents, 22% said such a failure happened more than once.
Among July respondents on the same base, 52% reported at least one such failure. The difference between the two months is too close to call. The share reporting exactly one such failure rose from 27% in July to 39% in August, a real change.
The question asks whether an organization had at least one such failure in a year, and whether it happened once or more than once. An enterprise that ships fifty agents has more chances to report a failure than an enterprise that ships two. The answers therefore count affected organizations and do not give a failure rate per agent.
Finding 3. Only 29% of enterprises with AI agents build production monitoring around real-time checks that alert when agent answers go wrong
Trace logging or gateway tracking, neither worded as checking whether an answer is correct, is the main monitoring approach for 53% of respondents whose organization deploys autonomous agents.
Production monitoring for AI agents can watch two different things. System health monitoring, such as trace logs and gateway metrics, shows whether an agent ran, how fast and at what cost. Quality monitoring checks whether the agent's answers are correct and can alert a team when quality drops. An agent can look healthy by every system measure while giving customers wrong answers, and system health monitoring will not flag the problem.
We asked respondents how their organization's live production monitoring for AI agents is set up: to evaluate system health, or to evaluate the quality of what the agent produces. We report the answers of respondents whose organization deploys autonomous agents.
Finding 3 — Only 29% of enterprises with AI agents build production monitoring around real-time checks that alert when agent answers go wrong
Base: 118 respondents whose organization deploys autonomous agents. One answer each. The 32% in the text leaves out respondents who did not know their organization's approach (Base: 105). The methodology gives the figures on all respondents.
Only 29% of respondents whose organization deploys autonomous agents named real-time automated quality checks on live agent output as the approach their production monitoring is built around. Among those who knew their organization's approach, the share is 32%.
The trace-logging option was worded as capturing infrastructure activity, token counts and raw inputs and outputs for later debugging. The gateway-tracking option was worded as tracking latency, error rates and cost. Neither option mentioned real-time evaluation of whether an answer is correct.
Trace logging or gateway tracking is the main approach for 53% of respondents whose organization deploys autonomous agents. Respondents could name only one approach, so some of them may also run quality checks.
Finding 4. 59% of August respondents run OpenAI's native evaluation tooling, against 31% of July respondents
OpenAI's native evaluations and traces are both the most widely used and the most often named primary platform.
An evaluation platform runs an agent against test cases, scores the results and keeps traces of each step the agent took. OpenAI builds evaluations and tracing into its developer platform, and Anthropic offers evaluation tools in its console. Independent tools such as Braintrust and Confident AI work alongside them. Without a platform, teams keep test cases and scores in scripts, which makes results harder to repeat and compare from one release to the next.
We asked respondents which agent reliability or evaluation platforms their organization uses today, and then which one is their primary platform.
Finding 4 — 59% of August respondents run OpenAI's native evaluation tooling, against 31% of July respondents
Base: 136 respondents who answered. Several answers allowed, so shares add up to more than 100%. Platforms named by fewer than 10 respondents are not shown. Base for the July figures in the text: 106.
The share of respondents whose organization runs OpenAI's native evaluation tooling was 59% in August, against 31% in July, which counts as a real change. The share running Confident AI's DeepEval went from 27% in July to 36% in August, a difference too close to call.
Finding 4 — Primary platform
Base: 136 respondents who answered. One answer each. Other platforms: Weights & Biases Weave 3%, custom in-house tooling 2%, LangSmith 2%, Arize / Phoenix 1%, Langfuse 1%, Promptfoo 1%, Datadog 1%. Bases for the primary-platform shares among users in the text: 80 for OpenAI's tooling and 49 for Confident AI.
Respondents named OpenAI's native evaluation tooling as their primary platform more often than any other: 40% did, against 12% for Braintrust and 10% each for Confident AI and Anthropic Claude Console / Workbench. Of the respondents whose organization runs OpenAI's tooling, 69% named it as primary. Of those running Confident AI, 29% named Confident AI as primary.
Finding 5. No reliability investment is named by more than 30% of enterprises as the one that will grow most next year, and the top two are too close to call
Human review workflows are named by 30% of respondents, production observability tooling by 26% and automated evaluation pipelines by 21%, while 11% say their reliability budget is not increasing.
Reliability budgets can go to people or to tooling. Human review workflows pay people to check agent output, which catches problems automated tests miss but costs more with every new agent and change. Observability tooling records what agents do, so teams can find problems after the fact. Automated evaluation pipelines grade output at scale, but only as well as their tests match real use.
We asked respondents which reliability or evaluation investment will grow most at their organization next year.
Finding 5 — No reliability investment is named by more than 30% of enterprises as the one that will grow most next year, and the top two are too close to call
Base: 140 respondents. One answer each.
Asked which single reliability investment will grow most next year, 30% of respondents named human review workflows, 26% named production observability tooling and 21% named automated evaluation pipelines. The gap between human review workflows and production observability tooling is too close to call.
In a separate question, 62% of respondents said their organization plans to adopt a new, additional or replacement agent reliability or evaluation platform within 12 months. Within three months, 35% expect to do so.
Finding 5 — Plans to adopt a new, additional or replacement platform
Base: 140 respondents. One answer each. The platforms-considered figures in the text allowed several answers (Base: 140).
Asked which platforms they are considering, 24% of respondents named Confident AI and 16% named Braintrust.
Finding 6. Only 9% of respondents chose "we trust automated evaluation today," and 91% named a limitation that reduces their trust in automated agent evaluations
Asked which limitation most reduces their trust in automated agent evaluations, 9% of respondents chose "we trust automated evaluation today," and 27% named poor alignment with real-world outcomes.
Automated evaluation uses scripts or another AI model to grade an agent's output, so a team can check far more cases than people could review by hand. The grades are useful only when they match what happens with real users. A team that does not trust its automated grades has to keep people reviewing output, which slows releases and costs reviewer time.
We asked respondents which limitation most reduces their trust in automated agent evaluations today. One of the answers respondents could choose was "we trust automated evaluation today."
Finding 6 — Only 9% of respondents chose “we trust automated evaluation today,” and 91% named a limitation that reduces their trust in automated agent evaluations
Base: 140 respondents. One answer each.
Only 9% of respondents chose "we trust automated evaluation today." Every other respondent named the one limitation that most reduces their trust, so their answers do not measure how much they trust automated evaluation overall. Poor alignment with real-world outcomes was chosen by 27% and lack of explainability by 24%.
Across all respondents, 2% of those reporting an evaluation-passing failure chose "we trust automated evaluation today," against 18% of those reporting no such failure. In July, 4% of respondents reporting a failure chose it, against 24% of those reporting none.
The two groups differ in more than failures. Nearly all respondents reporting a failure work at organizations that deploy autonomous agents, while 42% of the no-failure group work at organizations that do not.
Among respondents whose organization deploys autonomous agents, 3% of those reporting a failure chose the answer, against 10% of those reporting none, a gap too close to call. Among respondents reporting no failure, 29% at organizations that do not deploy agents chose the answer, against 10% at organizations that do, also too close to call. So we cannot tell from these answers whether having a failure changes which answer respondents choose.
Finding 6 — Chose “we trust automated evaluation today,” by group, August
Bases: 81, 50, 80, 29 and 21, in table order. Agent deployers are respondents whose organization deploys autonomous agents. July bases: 53 respondents reporting a failure and 41 reporting none.
What changed since the July Pulse
We asked the questions compared below in the same words in July and August. The bases for some rows are smaller, as the methodology explains.
What changed enough to call
Since the July Pulse — What changed enough to call, July → August
Changed enough to call means the significance test described in the methodology confirms the difference.
What did not change enough to call
Since the July Pulse — What did not change enough to call, July → August
Too close to call means the significance test described in the methodology does not confirm the difference, so we cannot tell whether the share moved.
What we did not compare
We did not compare primary evaluation platforms between the two months. In July, respondents could name a primary platform they had not listed as one they use. Among July respondents who named a specific platform as primary, 34% named one they had not listed.
We did not compare the platforms respondents are considering. July respondents could give only one answer to that question, and August respondents could give several.
We did not compare vendor selection factors or success metrics. In August we asked those questions only of respondents who named a specific platform as primary. In July we asked every respondent.
The bottom line: 56% of enterprises with AI agents let them ship production changes on automated checks alone or plan to, down from 75% in July
Among August respondents whose organization deploys autonomous agents, 42% expect to keep a human reviewing agents' production changes for the foreseeable future. In July, 20% of respondents in the same group said so.
Among respondents whose organization deploys autonomous agents, 56% already allow unreviewed production pushes for specific low-risk agents or changes, or are engineering toward them, against 75% in July.
Leaving out the 5% of August respondents whose organization runs no pre-deployment evaluations, 61% said their organization shipped an AI agent or LLM feature that passed its evaluations and then caused a customer-facing failure.
In the next Pulse, VentureBeat Intelligence will measure three things. The first is how the share of enterprises that expect to keep a human reviewing agents' production changes compares with August's 42%. The second is what share of respondents whose organization deploys autonomous agents name inline quality checks on live output as their monitoring approach, against August's 29%. The third is how the share running OpenAI's native evaluation tooling compares with August's 59%.
Respondent profile
Company size | Respondents |
100 to 499 employees | 37% (52) |
500 to 2,499 employees | 33% (46) |
2,500 to 9,999 employees | 15% (21) |
10,000 to 49,999 employees | 9% (12) |
50,000 or more employees | 6% (9) |
Base: 140 respondents. One answer each.
Purchasing role | Respondents |
Final decision-maker on AI purchasing | 53% (74) |
Technology recommender or influencer | 26% (36) |
End business user | 6% (8) |
No involvement in AI purchasing | 16% (22) |
Base: 140 respondents. One answer each.
Job role | Respondents |
Product or program manager | 20% (28) |
Director of data, AI or analytics | 11% (16) |
Director of engineering or IT | 11% (16) |
CIO, CTO or CISO | 10% (14) |
Consultant or advisor | 8% (11) |
VP of engineering, IT, data, AI or analytics | 8% (11) |
Investor or VC | 3% (4) |
Other roles | 29% (40) |
Base: 140 respondents. One answer each.
Industry | Respondents |
Healthcare or life sciences | 19% (27) |
Technology or software | 19% (26) |
Manufacturing or industrial | 15% (21) |
Retail or consumer | 13% (18) |
Financial services | 8% (11) |
Education | 7% (10) |
Other industries | 19% (27) |
Base: 140 respondents. One answer each.
Methodology
VentureBeat Intelligence fielded this VB Pulse survey in August and received 199 responses. Of those, 140 qualified for this report. Every figure is based on those 140 respondents unless a smaller number is stated with the figure. Some bases are smaller because a question applied to fewer respondents or some answers were left out; each table or the note under it gives its base.
Respondents are a self-selected group of VentureBeat readers and panel members. The figures describe these respondents and are approximate as a guide to the wider enterprise market. VentureBeat surveyed a separate sample each month, so month-to-month comparisons describe two different groups of respondents.
Across all respondents, including those whose organization does not deploy autonomous agents, 25% named inline quality checks as their monitoring approach, 31% of those who knew did, and 48% named trace logging or gateway tracking. The July comparison of inline quality checks in "What did not change enough to call" uses all respondents in both months.
A difference between two shares, whether between two groups in the same month or between July and August, counts as real, or enough to call, only when a statistical test gives p below 0.05; otherwise it is too close to call. We use a two-proportion z-test, or Fisher's exact test when either group has fewer than 40 respondents. To rank two answers within one question, we use an exact binomial test on single-answer questions and an exact McNemar test on questions that allowed several answers. Other rankings of answers within a question describe these respondents and are not tested.
Some groups in this report have fewer than 40 respondents: August agent deployers reporting no evaluation-passing failure, 29; August respondents at organizations not deploying agents and reporting no failure, 21; July agent deployers reporting no failure, 37. Percentages for groups this small are less precise; for a group of 21, a result could differ by about 21 percentage points either way from the true figure.
We compared many answers between the two months, so a few could pass the test by chance alone. The three changes behind the headline and Finding 4 have p-values of about 0.004 or lower.
The July report compared respondents with an evaluation-passing failure against respondents without one, on whether their organization already allows unreviewed production pushes or is building toward them. Recomputed among agent deployers, 87% of July respondents reporting a failure gave one of those answers, against 68% of those reporting none, a real difference. In August, the figures were 59% and 55%, too close to call.
The July report published the July figures for the production-push question on all 108 respondents: 67% already allowing unreviewed pushes or building toward them, and 18% expecting to keep a human reviewing agents' production changes. We use respondents whose organization deploys autonomous agents in both months. The July report also published the evaluation-passing failure figure on all 108 July respondents, at 49%. We leave out respondents whose organization runs no pre-deployment evaluations.
Among July respondents, 5% named OpenAI's tooling as their primary platform without listing it as one they use. Counting them as users, the July share is 36%, and the difference from August still counts as a real change.
In July, final purchasing decision-makers made up 44% of respondents, respondents at technology and software companies 14%, and respondents with no involvement in AI purchasing 24%. The respondent profile gives the August figures. None of the differences between the two months on these three measures is large enough to call.
