VentureBeat Intelligence surveyed respondents at organizations with 100 or more employees in August 2026. We asked where their AI agents get business context and whether those agents had given wrong answers traced to that context, once or more than once. We also asked how enterprises choose retrieval platforms.
Most enterprises have caught an AI agent giving a confidently wrong answer because of problems in their own company data. In the past six months, 64% of respondents traced at least one such answer to missing or inconsistent business context, such as a wrong metric definition or stale data. Of all respondents, 32% traced more than one. The question asked respondents to leave out errors by the model itself, and this report calls these answers context failures.
Enterprises use a governed semantic or context layer, which we call a semantic layer, to give AI agents and business intelligence (BI) tools one agreed understanding of company data. A semantic layer settles questions such as what counts as revenue or which customer table is current.
Enterprises that run, pilot or build a semantic layer were far more likely to report at least one context failure than enterprises only evaluating one or with no plans for one. Among those that run, pilot or build one, 78% reported a context failure, more than double the 37% among those only evaluating one or with no plans for one. We did not ask respondents why the two groups differ.
July Pulse respondents also showed a large gap. At enterprises that ran, piloted or built a semantic layer, 89% of respondents reported a context failure, against 35% at enterprises only evaluating one or with no plans for one.
Even enterprises that run a semantic layer seldom make it their agents' primary source of business context. Among enterprises running a semantic layer in production, only 23% name it as their agents' primary source. Across all respondents, retrieval over documents is named by 32% and direct queries to live systems by 21%, a gap too close to call.
Enterprises are also changing where their agents get context and which retrieval tools they run. Five answers changed enough to call since the July Pulse. More respondents named direct queries to live systems as their agents' main source of business context, 21% against 11% in July, and fewer named a mix that varies by use case.
The share of respondents running a custom in-house retrieval stack fell from 18% to 6%, while the share running the Qdrant vector database rose from 7% to 16%.
Among respondents who run a production retrieval system, 36% said ease of data ingestion most influenced their choice of retrieval infrastructure, up from 22% in July, a real change. Yet only 11% of respondents who run a production retrieval system named retrieval accuracy.
Model providers are bundling retrieval and memory into their own platforms, yet most respondents expect to keep at least some context tools independent of any one provider. The share who say their most likely direction is to consolidate onto a single model provider was 12% in July and 22% in August. The difference is too close to call.
Finding 1. 64% of enterprises have traced a confidently wrong AI agent answer to problems in their company data, and 49% of those more than once
In the past six months, 64% of respondents traced at least one confident but wrong agent answer to missing or inconsistent business context.
An AI agent answering a business question depends on company data and on the definitions behind it, such as what counts as revenue or which customer table is current. When a definition is missing, or two systems disagree, the agent can return a fluent answer built on the wrong figure. A confidently wrong answer is harder to catch than an obvious error, because nothing in the answer signals a problem. An employee who acts on the answer can pass a wrong number to a customer or make a decision on stale data.
We asked respondents whether, in the past six months, their AI agents had produced confident but wrong answers traced to missing or inconsistent business context. The question named wrong metric definitions and stale or misinterpreted data as examples, and asked respondents to exclude model error.
Finding 1 — Confident but wrong agent answers traced to business context, past six months
Base: 130 respondents. One answer each.
Of the respondents who reported a context failure, 49% said their agents had produced more than one.
Of all respondents, 31% said their agents had not produced a confident but wrong answer traced to business context.
Finding 2. Respondents at enterprises that run, pilot or build a semantic layer are far more likely to report a context failure than respondents at enterprises only evaluating one or with no plans for one
Among August respondents, 78% of those running, piloting or building a semantic layer reported a context failure, against 37% of those only evaluating one or with no plans for one. July's respondents showed a similar gap, and the gap holds when respondents who don't run agents on enterprise data or don't track root cause are left out.
A semantic layer holds one agreed definition for each business term and metric, and the agents and dashboards that query company data use those definitions. Without one, each agent or dashboard can apply its own definition, so two tools can give different answers to the same question. A semantic layer also gives a team a correct definition to check an agent's answer against. Building and governing one takes ongoing work, because someone has to agree each definition and keep it current.
We asked respondents whether their organization uses a governed semantic or context layer to give agents and BI tools a shared understanding of company data.
Finding 2 — Use of a governed semantic or context layer
Base: 130 respondents. One answer each.
Of all respondents, 67% said their organization runs a semantic layer in production or is piloting or building one.
Respondents to the July Pulse, a separate sample, also showed a large gap. In July, 89% of respondents at enterprises that ran, piloted or built a semantic layer reported a context failure. At enterprises only evaluating one or with no plans for one, 35% of July respondents did.
Finding 2 — Reported a context failure, by semantic-layer status
Base: 38 in the smaller August group and 34 in the smaller July group. Respondents who did not know whether their organization has a semantic layer are left out. Leaving out respondents who don't run agents on enterprise data or don't track root cause, the gap stays large in both months. August: 79% (Base: 86) against 41% (Base: 34). July: 90% (Base: 62) against 46% (Base: 26).
We did not ask respondents why these two groups differ, and we cannot separate the possible explanations. One is that teams with a semantic layer have a correct definition to compare a wrong answer against, so they find more of their context failures.
Another is that the enterprises that had context failures are the ones that went on to build a semantic layer.
Finding 3. Only 23% of enterprises running a semantic layer in production name it as their agents' primary source of business context
Retrieval over documents is named by 32% of respondents as their agents' primary source of business context, against 13% for a governed semantic layer.
Agents get business context in several ways. Retrieval pulls passages from documents or a vector index, direct queries read live systems such as databases, and long-context loading puts large inputs straight into the model. Each of those methods can leave an agent without the agreed definition of a business term, unless the agent also consults a semantic layer. An enterprise can run a semantic layer for its BI tools while its agents draw context from somewhere else.
We asked respondents what their AI agents use today as their primary source of business context when they need to understand enterprise data.
Finding 3 — Agents' primary source of business context
Base: 130 respondents. One answer each.
Finding 3 — Named a semantic layer as primary source, by group
Base: 48 respondents running a semantic layer in production and 87 that run, pilot or build one. Each row is a share of its own group.
Retrieval over documents or a vector index is the approach known as retrieval-augmented generation (RAG). Direct queries to live systems include SQL databases, APIs and Model Context Protocol (MCP) servers.
Respondents named RAG as their agents' primary source at 32% and direct queries to live systems at 21%, a gap too close to call. Loading large inputs directly into the model's context window was named by 19%.
Among all enterprises that run, pilot or build a semantic layer, 20% named it as their agents' primary source of business context.
Finding 4. Among respondents who run a production retrieval system, 36% said ease of data ingestion most influenced their choice of retrieval infrastructure, while 11% named retrieval accuracy
Among respondents running a production retrieval system, ease of data ingestion was named three times as often as retrieval accuracy as the factor that most influenced their choice.
Retrieval infrastructure stores a company's documents and data in searchable form and finds the passages an agent needs. Loading data in comes first, since an agent can find only what has been ingested, indexed and kept current. Retrieval accuracy describes whether the system returns the right passages. When retrieval hands an agent the wrong passage, the agent can give a confident answer built on it.
We asked respondents which factor most influenced their selection of retrieval (RAG) infrastructure. The figures below cover respondents who run a production retrieval system, all of whom answered the question.
Finding 4 — Factor that most influenced the choice of retrieval infrastructure
Base: 117 respondents who run a production retrieval system. One answer each.
Ease of data ingestion was named by 36% and access control and permissions by 22%, a gap too close to call.
The share naming ease of data ingestion rose from 22% in July to 36% in August, a real change among respondents who run a production retrieval system.
We can say which factor mattered most to each respondent. We cannot say how much weight respondents gave to accuracy.
Finding 5. Most enterprises expect to keep at least some retrieval and context tools independent of any model provider, while 22% expect to consolidate onto one model provider's stack
The share expecting to consolidate onto one model provider's native context stack was 12% in July and 22% in August, a difference too close to call.
Model providers now bundle retrieval, memory and orchestration into their own platforms. Consolidating onto one provider's stack means fewer tools to connect and one vendor to manage. The cost is dependence, since moving agents to another model later can mean rebuilding the context stack that feeds them. Standalone tools keep the context layer separate from the model, at the price of more integration work.
We told respondents that model providers are increasingly bundling retrieval, memory and orchestration directly into their platforms. We cited OpenAI's Codex plugins and connectors, Anthropic's Managed Agents memory and Google's Vertex AI Context Broker. We then asked for the most likely direction for their organization's retrieval and context stack over the next 12 months.
Finding 5 — Most likely direction over the next 12 months
Base: 130 respondents. One answer each. July base: 101.
Most respondents named a direction that keeps at least some of their context tools independent of any single model provider. Of all respondents, 37% said their most likely direction is to keep best-of-breed standalone tools.
Another 28% said a mix of provider-native and standalone tools by workload, and 2% said building their own context layer. Meanwhile, 22% said their most likely direction is to consolidate onto a single provider's native context stack.
Because the change from July is too close to call, we cannot say that more enterprises now expect to consolidate.
In a separate question, 65% of respondents said they plan to adopt a new, additional or replacement retrieval platform within 12 months. Within three months, 34% of respondents said they plan to adopt one.
Finding 6. Only 6% of enterprises now run a custom in-house retrieval stack, down from 18% in July, and nearly all enterprises with production retrieval run a system built by others
Among respondents who run a production retrieval system, 98% run at least one system they did not build, led by OpenAI retrieval and file search and Google Vertex AI Search.
A company can build its own retrieval stack or run a system built by a vendor or an open-source project. A custom stack gives a team full control over how documents are indexed and searched, but the team has to maintain it as models and data change. A system built by others can take less upkeep, but can tie the company's retrieval to that product's features, pricing and roadmap.
We asked respondents which retrieval systems they run in production.
Finding 6 — Retrieval systems in production
Base: 129 respondents who answered. Respondents could name several systems, so shares add up to more than 100%. July base: 101. Base for the 98% who run a system they did not build: 117 respondents who run a production retrieval system.
Finding 6 — Primary retrieval platform
Base: 74 respondents who named as primary a platform they also said they run. One answer each.
In production use, the gap between OpenAI retrieval and Google Vertex AI Search is too close to call. Google Vertex AI Search is ahead of Elasticsearch / OpenSearch, the next system at 19%, by enough to call.
Among respondents who named as primary a platform they also said they run, 35% named OpenAI retrieval and file search and 19% named Google Vertex AI Search, a gap too close to call. No other platform was named as primary by more than 10%.
The share of respondents running OpenAI retrieval was 46% in July and 45% in August. The share running Google Vertex AI Search was 41% in July and 34% in August. Both differences are too close to call.
The share of respondents running a custom in-house retrieval stack fell from 18% in July to 6% in August, a real change. The share running the Qdrant vector database rose from 7% to 16%, also a real change.
What changed since the July Pulse
We asked the questions compared below in the same words in July and August. Where a base is smaller than the full sample, because respondents skipped a question or only part of the sample applies, the table shows the smaller base.
What changed enough to call
What we measured | Base, July / August | July | August |
Name a mix that varies by use case as agents' primary source of business context | 100 / 130 | 17% (17) | 6% (8) |
Run a custom in-house retrieval stack in production | 101 / 129 | 18% (18) | 6% (8) |
Name direct queries to live systems as agents' primary source of business context | 100 / 130 | 11% (11) | 21% (27) |
Run Qdrant in production | 101 / 129 | 7% (7) | 16% (20) |
Name ease of data ingestion as the most influential selection factor, among respondents who run a production retrieval system | 94 / 117 | 22% (21) | 36% (42) |
Enough to call: the difference passes the significance test described in the methodology.
Fewer August respondents than July respondents named a mix that varies by use case as their agents' primary source of business context. Fewer also run a custom in-house retrieval stack.
More August respondents named direct queries to live systems as their agents' primary source, and more run Qdrant. Among respondents who run a production retrieval system, more named ease of data ingestion as the factor that most influenced their choice.
What did not change enough to call
What we measured | Base, July / August | July | August |
Expect to consolidate onto a single model provider's native context stack | 101 / 130 | 12% (12) | 22% (28) |
Reported at least one context failure in the past six months | 101 / 130 | 68% (69) | 64% (83) |
Reported context failures more than once | 101 / 130 | 37% (37) | 32% (41) |
Run a semantic layer in production | 101 / 130 | 32% (32) | 37% (48) |
Run OpenAI retrieval / file search in production | 101 / 129 | 46% (46) | 45% (58) |
Run Google Vertex AI Search in production | 101 / 129 | 41% (41) | 34% (44) |
Name response correctness as their primary success metric | 100 / 130 | 38% (38) | 28% (36) |
Expect tool-first or long-context retrieval, without a dedicated vector layer, to dominate by the end of 2026 | 101 / 130 | 15% (15) | 25% (32) |
Expect to keep best-of-breed standalone tools | 101 / 130 | 37% (37) | 37% (48) |
Plan to adopt a new, additional or replacement retrieval platform within 12 months | 100 / 130 | 52% (52) | 65% (84) |
Too close to call: the difference fails the significance test described in the methodology, so we cannot tell whether the share moved.
What we did not compare
We did not compare primary retrieval platforms between the two months. The question asked for one primary platform. In July, 18% of respondents gave several answers.
In August, 34% of respondents who answered the retrieval-systems question named a primary platform they had not said they run. Neither month's answers are clean enough to compare.
We also did not compare which retrieval providers respondents are considering. In August, every respondent with no plans to change platform also said they are not considering a change.
In July, only 37% of respondents with no plans to change also said they are not considering a change. The two answers line up in August and not in July, so we do not compare the question.
The mix of respondents in each month is shown below, and none of the differences between months is large enough to call.
Respondent group | July | August |
Final purchasing decision-makers | 38% (38) | 46% (60) |
Technology and software companies | 31% (31) | 22% (28) |
Organizations with more than 10,000 employees | 12% (12) | 5% (7) |
Base: 101 July respondents and 130 August respondents.
The bottom line: Most enterprises have traced wrong AI agent answers to their own company data, and those running, piloting or building a semantic layer are more likely to report one
Of all respondents, 64% reported a context failure in the past six months. Among enterprises that run, pilot or build a semantic layer, 78% reported one, against 37% of enterprises only evaluating one or with no plans for one.
Meanwhile, the share of respondents running a custom in-house retrieval stack fell from 18% in July to 6% in August. The share naming direct queries to live systems as agents' main context source rose from 11% to 21%. The share expecting to consolidate onto a single model provider's native stack was 12% in July and 22% in August, a difference too close to call.
In the next Pulse, VentureBeat Intelligence will measure three things. The first is whether the share of enterprises expecting to consolidate onto a model provider rises by enough to count as a real change. The second is whether more enterprises name the semantic layer as their agents' primary source of business context.
The third is whether enterprises that run, pilot or build a semantic layer continue to report context failures more often than enterprises only evaluating one or with no plans for one.
Respondent profile
Organization size | Respondents |
100 to 250 employees | 32% (42) |
251 to 1,000 employees | 34% (44) |
1,001 to 5,000 employees | 22% (28) |
5,001 to 10,000 employees | 7% (9) |
More than 10,000 employees | 5% (7) |
Base: 130 respondents. One answer each.
Role in AI purchasing | Respondents |
Final decision-maker on AI purchasing | 46% (60) |
Recommender or influencer on AI purchasing | 34% (44) |
User of AI solutions | 13% (17) |
No involvement in AI purchasing | 7% (9) |
Base: 130 respondents. One answer each.
Job level | Respondents |
Manager | 42% (55) |
VP or director | 22% (29) |
Individual contributor | 18% (24) |
C-suite | 8% (11) |
Other role, described in the respondent's own words | 8% (11) |
Base: 130 respondents. One answer each.
Industry | Respondents |
Technology or software | 22% (28) |
Healthcare or life sciences | 16% (21) |
Manufacturing | 15% (19) |
Financial services, banking or insurance | 12% (15) |
Education | 8% (11) |
Retail or e-commerce | 7% (9) |
Other industries | 18% (23) |
Industry described in the respondent's own words | 3% (4) |
Base: 130 respondents. One answer each.
Methodology
VentureBeat Intelligence fielded this VB Pulse survey in August 2026 and received 184 responses. Of those, 139 qualified for this report. Nine were removed because two of their answers about their own organization could not both be true, leaving 130. Every figure is based on those 130 respondents unless a smaller base is stated.
July figures use the bases published for the July 2026 Pulse, including the smaller bases the July report gave for individual questions.
Respondents are a self-selected group of VentureBeat readers and panel members, not a random sample of enterprises. Readers should not treat the figures as exact measures of the whole enterprise market. Respondents who described their role in their own words included reception, construction and social work, so some respondents work outside technology functions.
Some groups in this report have fewer than 40 respondents: those only evaluating a semantic layer or with no plans for one (August 38, July 34; 34 and 26 after the Finding 2 exclusions). Percentages for groups this small are less precise; for a group of 38, a result could differ by about 15 percentage points either way from the true figure.
A difference between two percentages, whether between two groups in the same month or between July and August, counts as real, or enough to call, when a statistical test gives p<0.05; otherwise it is too close to call. Two percentages from separate groups are compared with a two-proportion z-test, or with Fisher's exact test when either group has fewer than 40 respondents. Rankings between two answers from the same respondents use the exact McNemar test, which on a single-answer question is the exact binomial test.
