Enterprises are increasingly hesitant to give AI agents carte blanche ability to deploy code or system changes to production without human approval, even as they continue investing in tools designed to evaluate agent reliability.
In VentureBeat’s August Agentic Reliability and Evaluations wave, 56% of respondents whose organizations deploy autonomous AI agents said they already allow certain production changes based on automated evaluation results alone, without human review, or are engineering their systems to permit this within the next year. This is a notable drop from 75% in VentureBeat’s July survey.

One possible complication is respondent composition: More than half (53%) of August's 140 respondents were final AI purchasing decision-makers, compared with 44% in July.
However, the decline was also apparent within that group. Among final purchasing decision-makers whose organizations deploy autonomous agents, the share allowing unreviewed production changes or building toward them fell from 88% in July to 61% in August.
The findings therefore point to a meaningful shift within the surveyed organizations, rather than one explained simply by the different mix of respondents. Still, both surveys relied on separate, self-selected groups of VentureBeat readers and panel members, meaning the results do not necessarily establish a corresponding change across the entire enterprise market.
Interestingly, this movement was occurring before Anthropic CEO Dario Amodei published his explosive essay, "We Must Pace the Frontier" (on September 12), in which he called for slowing the pace of frontier AI development and strengthening safety oversight. Other industry leaders, including OpenAI CEO Sam Altman and Google DeepMind's Demis Hassabis, subsequently expressed support for greater caution.
Because the survey was conducted in August, the shift cannot be a direct reaction to Amodei's essay. However, the data does not establish what caused enterprises to become more cautious about autonomous deployments.
The more revealing divergence is between enterprises' willingness to trust automated evaluations with production deployment decisions and their continued adoption of the tools themselves. While support for eliminating human approval declined, use of OpenAI's native evaluation tooling increased substantially, from 31% in July to 59% in August.
Meanwhile, 61% of respondents whose organizations run pre-deployment evaluations reported at least one instance in which an AI agent or LLM feature passed internal testing but subsequently caused a customer-facing failure within the previous 12 months.

Together, the findings suggest enterprises are distinguishing between using automated evaluations to assess AI systems and relying on those evaluations as the sole authorization for putting changes into production.
The trust gap keeps human oversight central to enterprise AI
The August responses divide enterprises into distinct camps. More than a quarter (27%) already allow some low-risk agents to deploy changes without human validation, while 35% rule it out for the foreseeable future and 20% are engineering toward it within 12 months. Another 16% do not deploy autonomous agents at all. The remaining 2% were unsure. These figures cover all 140 August respondents, including organizations not deploying autonomous agents.
Excluding respondents whose organizations do not deploy autonomous agents, 32% already allow unreviewed production changes for specific low-risk agents or changes, while another 24% are building toward that capability. A further 42% expect to retain human review for the foreseeable future.
That last figure represents one of the survey's clearest changes: The share of respondents at agent-deploying organizations expecting to retain human approval rose from 20% in July to 42% in August, a statistically significant increase.
However, the shift toward retaining human approval does not appear to be accompanied by an equivalent shift in reliability spending priorities.

When asked which single reliability or evaluation investment would grow most over the next year, August respondents selected:
Human review workflows: 30%
Production observability tooling: 26%
Automated evaluation pipelines: 21%
Safety and policy evaluation: 11%
No increase in reliability budget: 11%
These figures measure which area respondents expect to experience the greatest investment growth, not how their organizations divide actual spending.
The July results were broadly similar: 31% selected human review workflows, 30% production observability, 19% automated evaluation pipelines and 16% safety and policy evaluation.
Human review therefore remained the most frequently selected growth area in August, although its lead over production observability was too small to establish a statistically meaningful difference. Nor do the results establish that enterprises are cutting evaluation spending or reallocating money away from automation.
Safety and policy evals go deeper than whether a model is accurate or not; they involve analyzing whether an agent will refuse when it should, adhere to policy, and avoid harmful outputs.
Although the share identifying safety and policy evaluation as the fastest-growing investment area declined from 16% to 11%, this does not mean actual spending on those activities fell. It indicates only that fewer August respondents selected that category as their organization's single largest expected area of reliability investment growth.
The survey instead paints a picture of enterprises continuing to develop both human and automated oversight, even as they reconsider which decisions to entrust to machines alone.
Internal evaluations still fail to prevent customer-facing incidents
One reason enterprises may be reluctant to eliminate human review is the persistent gap between internal evaluations and real-world performance. However, the survey does not directly establish that evaluation failures caused the change in deployment attitudes.
VB asked whether an AI or LLM feature that passed internal evals ultimately caused a customer-facing failure.
Leaving out the 5% of respondents with no pre-deployment evals, 61% reported at least one failure in the past 12 months. Staggeringly, 22% reported more than one such failure.
The directly comparable July reading was 53% of 101 respondents, rather than the previously published 49% of all 108 July respondents.
The August figure was numerically higher, but the difference was not statistically significant, meaning the survey does not establish that the proportion of organizations experiencing evaluation-passing failures actually increased.
One narrower finding did show a statistically significant change: The proportion reporting exactly one such failure rose from 27% in July to 39% in August.
The distinction matters because these results measure how many surveyed organizations experienced at least one failure, not how often individual agents or deployments fail. An organization deploying hundreds of AI features has more opportunities to encounter a customer-facing incident than one deploying only a handful.
Trust in systems meant to catch these issues is correspondingly thin: Just 9% of August respondents selected the answer "we trust automated evaluation today" when asked which limitation most reduces their trust in automated agent evaluations. This is compared with 13% in July and 5% in June.
The decrease from July to August was not statistically significant. Moreover, because respondents were asked to select their single biggest concern, the finding should not be interpreted as meaning that 91% have no confidence whatsoever in automated evaluations.
The biggest objection is that evaluation systems do not align with real-world outcomes (identified by 27% of respondents). Other concerns are lack of explainability (24% of respondents), evaluation bias or inconsistency (16%), data leakage or privacy concerns (13%), and tooling immaturity (12%).
Interestingly, respondents who had experienced evaluation-passing failures were not meaningfully less inclined to pursue autonomous deployments than those who had not.
Among August respondents whose organizations deploy autonomous agents, 59% of those reporting an evaluation-passing failure already permit unreviewed production changes or are engineering toward them. The comparable figure was 55% among respondents reporting no such failure, a difference too small to establish a meaningful relationship.
This is a notable contrast with the July survey, in which organizations that had experienced evaluation-passing failures appeared more likely to be pursuing zero-human deployment. That relationship did not hold up clearly in August.
Ultimately, the August results support a more nuanced conclusion than enterprises simply abandoning automation because automated evaluations do not work. Organizations are becoming more reluctant to remove human approval from production deployments, while continuing to invest in and adopt evaluation infrastructure.
The emerging question is not whether enterprises will automate agent evaluation, but how much operational authority they will allow those evaluations to exercise without human intervention.
How enterprises are testing agents post-deployment
But pre-deployment testing is only one layer of assurance. Just as importantly, enterprises need to understand what an agent does after it goes live.
Among the 118 August respondents whose organizations deploy autonomous agents, only 29% said their primary production monitoring approach uses real-time automated quality checks on live agent outputs.
By comparison, 36% primarily rely on transaction trace logging, which records infrastructure activity, token counts and inputs and outputs for subsequent debugging, without necessarily evaluating whether the agent's answer was correct.
This is troubling, because a fast, confidently wrong response can produce a clean trace and a 200 status code, meaning an HTTP request was successfully completed.
The full breakdown of primary production monitoring approaches among organizations deploying autonomous agents was:
Transaction trace logging: 36%
Inline quality assertions on live traffic: 29%
API gateway infrastructure tracking: 16%
Don't know or not my area: 11%
Ad-hoc review: 8%
Combined, 53% named either transaction tracing or gateway monitoring as their primary approach. Both can help identify infrastructure problems, but neither was described in the questionnaire as directly checking the correctness of agent-generated outputs.
The monitoring gap becomes particularly apparent among organizations already permitting autonomous production deployments. According to VentureBeat Intelligence's underlying August survey analysis, just 10 of the 38 respondents whose organizations already allow certain production changes without human approval identified automated quality checks on live outputs as their primary monitoring approach.
That amounts to approximately 26%, similar to the 28% recorded among 40 comparable respondents in July. The findings suggest that removing human approval from deployment does not necessarily mean enterprises are making automated output-quality monitoring their primary safeguard in production.
However, the results do not establish that the remaining organizations lack other quality controls. Some may run supplementary automated checks or rely on additional forms of review that were not captured by the single-choice question.
Nevertheless, the results suggest that production monitoring at many enterprises remains more heavily oriented toward determining whether an AI system is functioning than whether it is generating correct answers.
That distinction is particularly significant when organizations permit AI agents to deploy changes without a human approval step.
Eliminating human review at deployment does not necessarily eliminate automated mechanisms for detecting problems afterward. But when an organization relies primarily on infrastructure monitoring, a plausible-looking yet incorrect agent output may not generate an error or alert.
The August survey does not establish how many enterprises allowing unreviewed deployments lack live quality monitoring altogether, nor whether their problems are typically discovered through automated alerts, employees or customer complaints.
It does reveal a practical challenge for enterprises seeking greater agent autonomy: Automated evaluations must catch more than infrastructure failures. They must also identify outputs that are syntactically valid, delivered successfully and nevertheless wrong.
For organizations expanding their use of agents, a critical question is how to establish effective quality monitoring after deployment without relying exclusively on customers or human reviewers to discover problems.
The vendor market
AI evaluation and observability tooling is similarly mixed.
OpenAI's developer platform appears in 59% of stacks (compared to 31% in July), a statistically significant increase, and was named the primary evaluation platform by 40% of August respondents who answered the platform question.

Among organizations specifically using OpenAI's native evaluation tooling, approximately 69% identified it as their primary platform. These are two different measures: The first describes OpenAI's primary-platform share across all respondents, while the second measures how often organizations already using OpenAI select it as their main platform.
Confident AI is in 36% of stacks, with 10% of all respondents naming it as their primary platform. Among organizations using Confident AI, 29% identified it as their primary platform.
Meanwhile, Braintrust is in 23% of stacks and serves as a primary platform for 12% of all respondents; Anthropic sits in 16% of stacks, serving as primary for 10% of all respondents; and LangSmith is in 13% of stacks, serving as primary for 2% of all respondents.
For consistency, the comparison below uses the same base of 136 August respondents for both platform adoption and primary-platform selection.
Platform | Used in stack | Primary platform |
OpenAI Developer Platform | 59% | 40% |
Confident AI (DeepEval) | 36% | 10% |
Braintrust | 23% | 12% |
Anthropic Claude Console / Workbench | 16% | 10% |
LangSmith | 13% | 2% |
Respondents could name multiple tools in use but only one primary platform, so the adoption percentages are not mutually exclusive.
The strongest month-over-month vendor finding is OpenAI's increased adoption, which rose from 31% to 59%. Confident AI also rose from 27% in July to 36% in August, although that difference was not statistically significant.
VentureBeat Intelligence did not compare vendors' primary-platform shares between the two months because the July survey permitted respondents to select a primary platform they had not identified as one they used, complicating a like-for-like comparison.
As for infrastructure roadmaps, 62% of August respondents plan to adopt a new, additional or replacement evaluation platform within 12 months. More than a third (35%) expect to do so within three months.
These plans do not necessarily mean organizations are dissatisfied with or abandoning their existing evaluation vendors. The survey includes enterprises seeking to add another platform alongside their current tools, as well as those contemplating replacements.
For vendors, this signals a gap between serving as a primary driver in a tech stack, or merely existing within them. As technologies mature, this could create a different competitive position and renewal risk.
However, the survey does not directly measure vendor retention, renewal rates or whether organizations using a platform as a secondary tool are more likely to discontinue it. Those remain potential competitive implications rather than established outcomes.
The enterprise AI trust gap is evolving, not disappearing
Taken together, VentureBeat Intelligence's August findings expose a tension in the enterprise AI market: Organizations continue to adopt evaluation infrastructure even as they grow more cautious about relying on automated evaluation results to authorize production changes without human approval.
The percentage of respondents at organizations deploying autonomous agents who already permit unreviewed deployment or are building toward it fell significantly, from 75% to 56%. The share expecting to retain human review for the foreseeable future more than doubled, from 20% to 42%.
At the same time, most respondents whose organizations run pre-deployment evaluations reported at least one customer-facing failure after a feature passed internal testing, while only 29% of respondents at organizations deploying autonomous agents named real-time output-quality checks as their primary monitoring approach.
Those findings suggest unresolved gaps in the systems enterprises rely on to judge agent reliability. They do not, however, prove that evaluations are becoming less effective or that the reported failures directly caused organizations to rethink deployment autonomy.
And enterprises are hardly giving up on the technology: OpenAI's evaluation tooling gained substantial adoption, while nearly two-thirds of August respondents expect to adopt, add or replace an evaluation platform within a year.
The message for enterprise AI teams is that automating evaluation and automating deployment approval are different decisions, with different consequences when the underlying tests fail.
As agents gain the ability to make changes to increasingly important business systems, enterprises must decide not simply whether automated evaluations are useful, but when their results are reliable enough to justify removing a human from the final approval process.
The August survey suggests that, for a growing share of its respondents, that threshold has not yet been met.
