VentureBeat Intelligence surveyed respondents at organizations with 100 or more employees in August. We asked how they choose, run and judge their AI infrastructure. Among respondents running AI in production, Google's Gemini and Google Cloud together are named as the primary AI platform more than twice as often as OpenAI.

Google is named as the primary AI platform more often than any other company. Among respondents running AI in production whose primary platform is one they use, 48% name a Google platform, Gemini or Google Cloud, as their primary AI platform. OpenAI is the primary platform for 22% and Microsoft Azure for 13%.

Google leads because its share combines two platforms near the top of their lists, Gemini among models and Google Cloud among clouds. Among respondents running AI in production, 53% use Gemini and 56% use OpenAI, a gap too close to call. Google Cloud is used by 34% and Microsoft Azure by 37%, also a gap too close to call.

Counting any use, 65% of respondents running AI in production use at least one Google platform. OpenAI is used by 56% and Azure by 37%. Google's lead over Azure on use is large enough to call. The gap between Google and OpenAI on use is too close to call.

Claims that AI infrastructure spending is measured rigorously come mostly from the people with final say on AI purchasing. Final decision makers make up 62% of the respondents who say return is tracked rigorously.

Among final decision makers on AI purchasing, 64% say their organization tracks the return on its AI infrastructure spending rigorously, against 34% of all other respondents. July's respondents showed a similar gap.

Among respondents running AI in production, those who say they track return rigorously also rate their platform as better value. Rigorous trackers give value for money an average of 4.24 out of 5, against 3.50 from the rest. We cannot tell whether tracking return raises the rating, or whether satisfied buyers are quicker to call their tracking rigorous.

Among respondents running AI in production who gave a figure, 80% put average accelerator utilization at 50% or below.

Few respondents name unit price as a buying factor. Only 11% of respondents name cost per 1M tokens as a factor in their choice, while 40% name integration with the cloud or data stack they already run.

Asked which emerging compute option their organization is most likely to evaluate in the next 12 months, 35% of respondents named specialized AI clouds such as CoreWeave, Lambda, Crusoe and Nebius. Non-Nvidia accelerators are named by 26%.

Finding 1. 48% of respondents running AI in production name a Google platform as primary, more than twice OpenAI's share

Google is the only company with a platform near the top of both the model list and the cloud list, while Gemini and Google Cloud are each too close to call against their closest rival.

A primary AI platform is the service an enterprise relies on first to run its AI work, whether a model provider, a cloud or its own infrastructure. Some companies sell only AI models, while the big cloud providers also offer models of their own. Among the cloud providers, the survey listed a model option only for Google. Buying models and cloud from one company can simplify contracts and data handling, but it ties more of the AI stack to one supplier.

We asked which AI compute platforms the organization uses in production, allowing several answers, and which one is its primary platform. We report the answers of respondents whose organizations run AI in production, at scale or for some workloads.

Finding 1 — 48% of respondents running AI in production name a Google platform as primary, more than twice OpenAI's share

Google (Gemini or Google Cloud)
48%
Primary platform
65%
Use in production
OpenAI
22%
Primary platform
56%
Use in production
Microsoft (Azure)
13%
Primary platform
37%
Use in production
Oracle (Oracle Cloud Infrastructure)
7%
Primary platform
18%
Use in production
Amazon (AWS)
4%
Primary platform
13%
Use in production
Anthropic
4%
Primary platform
16%
Use in production
Custom open-source self-managed stack
2%
Primary platform
9%
Use in production

Primary platform: base 54 respondents running AI in production whose primary platform is one they listed as in production use, one answer each. Google combines Gemini at 28% and Google Cloud at 20%. The methodology gives the figures on all 68.

Use in production: base 68 respondents whose organizations run AI in production. Several answers allowed, so shares add up to more than 100%.

Among respondents running AI in production whose primary platform is one they use, 48% name Gemini or Google Cloud as their primary platform. OpenAI is primary for 22% and Microsoft Azure for 13%. Google's share is more than twice OpenAI's and nearly four times Azure's, and both gaps are large enough to call.

On use, 65% of respondents running AI in production use at least one Google platform, against 56% for OpenAI and 37% for Azure. Google's lead over Azure is large enough to call, but the gap between Google and OpenAI on use is too close to call.

Finding 1 — Model provider in production use

56%
OpenAI
53%
Google (Gemini)
16%
Anthropic

Base: 68 respondents whose organizations run AI in production. Several answers allowed, so shares add up to more than 100%.

Finding 1 — Cloud provider in production use

37%
Microsoft Azure
34%
Google Cloud (GCP)
18%
Oracle Cloud Infrastructure
13%
AWS

Base: 68 respondents whose organizations run AI in production. Several answers allowed, so shares add up to more than 100%. Specialized AI clouds such as CoreWeave, Lambda, Crusoe and Nebius are used by 9% and are not shown.

Also not shown: Baseten, Anyscale, Fireworks or Together at 9%, custom open-source self-managed stacks at 9% and on-prem or colocated GPU clusters at 4%.

Google leads because its share combines two platforms near the top of their lists, and no other company has a platform near the top of both. On models, 53% of respondents running AI in production use Gemini and 56% use OpenAI. On clouds, 34% use Google Cloud and 37% use Azure. Both gaps are too close to call.

As a primary platform, Gemini is too close to call against OpenAI, and Google Cloud is too close to call against Azure. Among respondents whose primary platform is one they use, Gemini is primary for 28% and OpenAI for 22%. Google Cloud is primary for 20% and Azure for 13%.

Finding 2. Final decision makers are nearly twice as likely as other respondents to say the return on AI infrastructure spending is tracked rigorously

Among final decision makers on AI purchasing, 64% say their organization tracks the return on AI infrastructure spending rigorously, against 34% of all other respondents.

Tracking the return on AI infrastructure means tying what an organization spends on compute to what that spending produces, such as revenue, savings or faster work. Without that link, a team has a harder time telling which workloads earn their cost and which to cut. Budget decisions then rest on estimates.

We asked whether the organization can quantify the return on its AI infrastructure spending today. Across all respondents, 48% say they track it rigorously, 38% say partially and 15% say they cannot quantify it yet.

Finding 2 — Final decision makers are nearly twice as likely as other respondents to say the return on AI infrastructure spending is tracked rigorously

48%
Yes, we track it rigorously
38%
Partially
15%
No, we can't quantify it yet

Base: 109 respondents. One answer each.

Finding 2 — Say return is tracked rigorously

64%
Final decision makers on AI purchasing, August (50 respondents)
34%
All other respondents, August (59 respondents)
62%
Final decision makers on AI purchasing, July (61 respondents)
35%
All other respondents, July (97 respondents)
71%
Respondents running AI at scale, August (28 respondents)
40%
Respondents not running AI at scale, August (81 respondents)
61%
Final decision makers not running AI at scale, August (33 respondents)
25%
Others not running AI at scale, August (48 respondents)

Bases: August 109 respondents; July 158 respondents who gave one answer to both questions. Group sizes are in the last column.

Final decision makers on AI purchasing say their organization tracks return rigorously at nearly twice the rate of everyone else. July's respondents showed a similar gap: 62% of final decision makers against 35% of others.

Respondents at organizations running AI at scale are more likely to claim rigorous tracking, at 71% against 40% of the rest. The gap between final decision makers and others does not come from respondents running AI at scale alone.

Among respondents whose organizations are not yet running AI at scale, 61% of final decision makers claim rigorous tracking, against 25% of others.

Most of the claims of rigorous tracking come from final decision makers. Final decision makers make up 46% of all respondents but 62% of the respondents who say return is tracked rigorously.

We cannot tell why the two groups answer differently. Final decision makers may see figures the rest of the organization does not. Final decision makers may also be grading a purchase they made.

Finding 3. Production respondents who say they track return rigorously rate their platform as better value

Among respondents running AI in production, those who say return is tracked rigorously rate their primary platform's value for money at 4.24 out of 5, against 3.50 from the rest.

Value for money is the rating where a buyer weighs a platform's price against what it delivers. The rating depends on what the buyer can see of both, including any figures on the return the spending produces.

We asked respondents to rate their primary AI compute platform from 1 (very poor) to 5 (excellent) on overall satisfaction, ease of implementation and value for money. We report the ratings of respondents whose organizations run AI in production, at scale or for some workloads.

Finding 3 — Production respondents who say they track return rigorously rate their platform as better value

Overall satisfaction
4.50
Track return rigorously (34)
3.79
Everyone else (34)
4.15
All production respondents (68)
Ease of implementation
4.35
Track return rigorously (34)
3.76
Everyone else (34)
4.06
All production respondents (68)
Value for money
4.24
Track return rigorously (34)
3.50
Everyone else (34)
3.87
All production respondents (68)

Base: 68 respondents whose organizations run AI in production; group sizes in parentheses. Average scores on a scale of 1 (very poor) to 5 (excellent).

Across production respondents, value for money averages 3.87, against 4.06 for ease of implementation and 4.15 for overall satisfaction. The gap between the two groups is about three quarters of a point on value for money and on overall satisfaction. On ease of implementation the gap is about 0.6 points. In all, 85% of rigorous trackers score value for money a 4 or 5, against 56% of the others.

July's respondents showed the same direction on value for money, overall satisfaction and ease of implementation. Across all of July's respondents, rigorous trackers gave value for money 4.22 and everyone else gave 3.62.

We cannot tell which way the relationship runs. Measuring return may help a team see what the spending buys. A team already pleased with its spending may also be quicker to call its tracking rigorous. Finding 2 shows who makes the rigorous claim most often.

Finding 4. 80% of respondents running AI in production who gave a figure put accelerator utilization at 50% or below

Among respondents running AI in production who gave a utilization figure, 80% put average GPU or accelerator utilization at 50% or below, and 40% put it at 25% or below.

Accelerator utilization is the share of time an organization's GPUs or other AI chips spend doing useful work. Owned or reserved accelerators cost money whether they run or sit idle, so low utilization on that capacity means paying for chips that produce nothing. Utilization can fall when capacity is reserved for peak demand.

We asked for the organization's average GPU or accelerator utilization in production. We report the answers of respondents whose organizations run AI in production, at scale or for some workloads.

Finding 4 — 80% of respondents running AI in production who gave a figure put accelerator utilization at 50% or below

7%
Under 10%
31%
10–25%
38%
26–50%
19%
Over 50%
3%
We don't measure it
1%
We don't operate our own GPUs, we consume via API

Base: 68 respondents whose organizations run AI in production. One answer each. Base for shares in the text: 65 who gave a figure, or 45 without respondents listing only model-API providers.

Only 20% of the production respondents who gave a figure report utilization above half. Leaving out respondents who list only model-API providers as their production platforms, 78% of the rest report utilization at half or below, so the result does not depend on that group.

Respondents running AI at scale and those with some workloads in production differ on utilization by too little to call. Among those giving a figure, 29% of respondents running AI at scale report utilization above half, against 14% of the rest.

Telemetry studies use a different measure. Cast AI's 2026 State of Kubernetes Optimization report measured average GPU utilization of 5% across the Kubernetes clusters it analyzed. Cast AI's figure covers different organizations, so the two figures are not comparable.

Finding 5. Only 11% of respondents name cost per 1M tokens as a factor in choosing AI infrastructure

Integration with the existing cloud or data stack is named by 40% of respondents as a buying factor, while 11% name cost per 1M tokens.

Cost per 1M tokens is the price a provider charges to process a million tokens, the units of text a model reads and writes. The figure makes providers easy to compare, but the total bill also depends on how many tokens a workload uses and on the infrastructure around it. Integration with an existing cloud or data stack decides how much of that surrounding work a team must do.

We asked which factors most influenced the organization's choice of AI infrastructure, allowing several answers.

Finding 5 — Only 11% of respondents name cost per 1M tokens as a factor in choosing AI infrastructure

40%
Integration with existing cloud or data stack
37%
Performance (latency, throughput)
34%
Security or compliance requirements
26%
Total cost of ownership
24%
Fine-grained autoscaling for spiky workloads
11%
Cost per 1M tokens

Base: 109 respondents. Several answers allowed, so shares add up to more than 100%. A seventh option is not reported; the methodology explains why.

Integration with the existing stack is named by 40%, performance by 37% and security or compliance requirements by 34%. Total cost of ownership is named by 26%.

Cost per 1M tokens is named by 11% of respondents, against 40% for integration with the existing stack.

Finding 6. Uptime and developer productivity lead the metrics used to judge AI infrastructure, ahead of any AI-specific measure

Uptime is used by 56% and developer productivity by 46%, ahead of latency at 28%, the most-used AI-specific metric.

The metrics a team tracks decide what problems it notices. Uptime and developer productivity show whether systems stay up and teams ship quickly. AI-specific metrics, such as time to first token, tokens per second and cost per 1M tokens, show how fast a model starts answering, how much it produces and what each answer costs. A team that tracks only general metrics can miss an AI workload that is running but slow or expensive.

We asked which metrics the organization uses to judge whether its AI infrastructure is working, allowing several answers.

Finding 6 — Uptime and developer productivity lead the metrics used to judge AI infrastructure, ahead of any AI-specific measure

56%
Uptime or reliability
46%
Developer productivity or deployment speed
28%
Latency (time to first token)
27%
Throughput (tokens per second, QPS)
23%
Cost per 1M tokens

Base: 108 respondents who answered. Several answers allowed, so shares add up to more than 100%. Two write-in answers are not shown.

Uptime and developer productivity are each used by more respondents than latency, at 28%. Neither uptime nor developer productivity is specific to AI, and a team could apply both to a database or a build system.

Among respondents who track cost per 1M tokens, 44% say they track return rigorously, against 49% of those who do not track token cost. The gap is too close to call.

Finding 7. 35% of respondents are most likely to evaluate specialized AI clouds next among emerging compute options

Specialized AI clouds such as CoreWeave, Lambda, Crusoe and Nebius are named by 35% for evaluation in the next 12 months, and non-Nvidia accelerators by 26%.

Specialized AI clouds, such as CoreWeave, Lambda, Crusoe and Nebius, rent out GPU capacity built for AI work, alongside or instead of the big cloud providers. Non-Nvidia accelerators, such as Google's TPU and AWS Trainium, offer another source of AI chips, mostly through the big clouds. Evaluating either can widen an organization's options on price and supply.

We asked which emerging compute option the organization is most likely to evaluate in the next 12 months, asking for one answer.

Finding 7 — 35% of respondents are most likely to evaluate specialized AI clouds next among emerging compute options

35%
AI-specialized clouds (CoreWeave, Lambda, Crusoe, Nebius)
26%
Non-Nvidia accelerators (Trainium, TPU, AMD Instinct, Gaudi, in-house ASICs)
19%
Nvidia Blackwell (GB300) or next-generation GPUs
9%
None of the above
8%
Decentralized or distributed compute networks
3%
Sovereign or region-specific compute

Base: 109 respondents. One answer each.

Finding 7 — Expected use over the next 12 months

Specialized AI clouds (100 respondents)
52%
Doing more
7%
Doing less
Inference APIs (87 respondents)
40%
Doing more
9%
Doing less
On-prem or colocated (93 respondents)
38%
Doing more
12%
Doing less
Hyperscalers (97 respondents)
35%
Doing more
11%
Doing less

Bases are respondents who gave an answer on the scale for each item; the rest answered about the same. The specialized-cloud item named no providers.

Specialized AI clouds are the only approach where a majority of the respondents who answered expect to do more. Of those respondents, 52% expect to do more with specialized AI clouds over the next 12 months.

Most respondents plan to add or replace a platform within a year. In all, 61% plan to adopt a new, additional or replacement platform in the next 12 months, and 34% plan to do so within three months.

Finding 7 — Plan to adopt a new, additional or replacement platform

34%
Yes, within 0–3 months
23%
Yes, within 3–6 months
4%
Yes, within 6–12 months
39%
No plans to change

Base: 109 respondents. One answer each.

Finding 8. Dell is named by 38% of respondents for inference memory and Nvidia by 31%, while 17% name no approach at all

Dell is named by 38% of respondents and Nvidia by 31%, while 17% name no approach and say they have not addressed inference-memory limits or are not aware of this as a constraint.

When a model answers a long prompt or a long conversation, it keeps intermediate results, called the KV cache, in memory so it does not recompute them for each new token. The cache grows with the length of the context and the number of users, and it competes with the model for scarce GPU memory. Vendors now sell storage and software that move the cache off the GPU.

We asked which inference-memory or KV-cache approaches or vendors the organization is using or evaluating, allowing several answers. The question opened by stating that the binding constraint on inference is shifting from GPU compute to memory. The answers respond to that framing.

Finding 8 — Dell is named by 38% of respondents for inference memory and Nvidia by 31%, while 17% name no approach at all

38%
Dell (PowerScale, Project Lightning)
31%
Nvidia (Dynamo, ICMSP)
20%
Open-source KV-cache tooling (LMCache, vLLM prefix caching)
17%
Model-level efficiency (MLA, quantization)
17%
Hammerspace (Tier Zero)
17%
VAST Data (Undivided Attention)
11%
DDN (Infinia)
9%
WEKA (Augmented Memory Grid)
13%
Have not addressed inference-memory limits yet
8%
Not aware of this as a constraint

Base: 109 respondents. Several answers allowed, so shares add up to more than 100%.

Open-source KV-cache tooling is named by 20% and model-level efficiency by 17%. Hammerspace and VAST Data are each named by 17%, DDN by 11% and WEKA by 9%.

In all, 20% of respondents say they have not addressed inference-memory limits yet or are not aware of this as a constraint. Of those respondents, 14% also named an approach they use or are evaluating. Leaving them out, 17% of respondents name no approach.

What changed since the July Pulse

The July Pulse asked most of the same questions in the same words. Five measures can be compared, and one changed by enough to call.

What changed enough to call

What changed enough to call

22% → 12%
Expect to do less with on-prem or colocated compute

Bases: respondents who gave an answer on the scale for on-prem or colocated compute, 149 in July and 93 in August.

Fewer respondents expect to cut back on on-premises or colocated compute. The share fell from 22% in July to 12% in August.

Fewer technology companies and fewer organizations above 1,000 employees answered in August than in July. Part of the change may reflect who answered.

What did not change enough to call

What did not change enough to call

46% → 48%
Say return is tracked rigorously
76% → 80%
Production utilization at 50% or below, among those giving a figure
60% → 61%
Plan to adopt a new, additional or replacement platform within 12 months
42% → 52%
Expect to do more with specialized AI clouds

Too close to call: none of these differences passes the test described in the methodology.

What we did not compare

Four questions changed form between July and August. In July, the form capped the buying-factor and success-metric questions at two answers, and in August it did not. The July form took one answer on inference memory, and the August form took several. The July form also accepted several answers on the emerging-option question, and many July respondents gave several, so we do not compare that question either.

We do not compare the average value-for-money rating with July. July's figure covers all of July's respondents, while August's ratings cover only those running AI in production.

We do not compare platform use or primary platform with July, including Google, OpenAI and Azure as primary platform and use of specialized AI clouds. July's platform figures cover all of July's respondents, while August's cover only those running AI in production, so the bases differ.

The bottom line: Google is the primary AI platform for 48% of respondents running AI in production, more than twice OpenAI's share

Among respondents running AI in production whose primary platform is one they use, 48% name Gemini or Google Cloud as their primary platform. OpenAI is primary for 22% and Microsoft Azure for 13%.

Google leads because its share combines two platforms near the top of their lists, Gemini among models and Google Cloud among clouds. On use, Gemini against OpenAI and Google Cloud against Azure are both too close to call. We did not ask respondents why they chose their primary platform.

Two other findings bear on how organizations judge AI infrastructure spending. Among respondents running AI in production who gave a figure, 80% put average accelerator utilization at 50% or below. Final decision makers on AI purchasing are nearly twice as likely as other respondents to say return is tracked rigorously.

Respondent profile

Organization size

Respondents

100–250 employees

32% (35)

251–1,000 employees

29% (32)

1,001–5,000 employees

21% (23)

5,001–10,000 employees

6% (6)

10,001 or more employees

12% (13)

Base: 109 respondents. One answer each.

AI deployment stage

Respondents

AI in production at scale

26% (28)

Some workloads in production

37% (40)

Experimenting or proofs of concept

33% (36)

Not yet running AI workloads

5% (5)

Base: 109 respondents. One answer each.

Role in AI purchasing

Respondents

Final decision maker

46% (50)

Recommender or influencer

27% (29)

User

17% (18)

No involvement

11% (12)

Base: 109 respondents. One answer each.

Job level

Respondents

Manager

47% (51)

Individual contributor

22% (24)

VP or director

15% (16)

C-suite

14% (15)

Other role (write-in)

3% (3)

Base: 109 respondents. One answer each.

Industry

Respondents

Technology or software

17% (18)

Healthcare or life sciences

17% (18)

Manufacturing

16% (17)

Retail or e-commerce

11% (12)

Financial services, banking or insurance

9% (10)

Transportation or logistics

7% (8)

Education

7% (8)

Government or public sector

4% (4)

Telecommunications or media

3% (3)

Energy or utilities

3% (3)

Professional services or consulting

2% (2)

Other industries (write-in)

6% (6)

Base: 109 respondents. One answer each. Write-ins were coded where clear: "Retail" and "Walmart" to retail, "Nurse" to healthcare and "Transportation" to transportation.

Methodology

VentureBeat Intelligence fielded this VB Pulse survey in August and received 179 responses. Of those, 118 came from organizations with 100 or more employees. Nine were removed because they said their organization is not yet running AI workloads and also named AI platforms it uses in production, two answers that cannot both be true, leaving 109. Every figure is on those 109 unless a smaller base is stated with the figure.

Respondents are a self-selected group of VentureBeat readers and panel members, not a probability sample. The figures describe these respondents and apply to the wider market only as far as their mix allows. The base includes write-in roles outside technology and management, such as daycare teaching and nursing.

Respondents who are experimenting or running proofs of concept also ticked platforms their organization uses in production. Platform figures, platform ratings and utilization figures therefore cover only the 68 respondents running AI in production, at scale or for some workloads.

Company figures count a respondent once for each company, so a respondent who uses both Gemini and Google Cloud counts once for Google. The survey listed OpenAI and Microsoft Azure as separate options and did not ask whether respondents reach OpenAI's models through Azure.

Every respondent running AI in production named a primary platform, but 21% named one they had not listed as in use. Counting all 68, including those respondents, Google is primary for 47%, OpenAI for 22% and Azure for 13%.

A seventh buying-factor option, access to GPUs or availability, is not reported because we could not confirm that every respondent saw it.

Some groups in this report have fewer than 40 respondents: respondents running AI at scale, 28; final decision makers not running AI at scale, 33; respondents running AI in production who track return rigorously, 34, and the other 34; respondents with some workloads in production who gave a utilization figure, 37; respondents who track cost per 1M tokens, 25; and respondents who have not addressed inference-memory limits or are not aware of this as a constraint, 22. Percentages for groups this small are less precise; for a group of 34, a result could differ by about 17 percentage points either way from the true figure.

A difference between two groups, or a change between July and August, counts as enough to call only when a statistical test gives p<0.05; otherwise it is too close to call. Two percentages from separate groups are compared with the two-proportion z-test, or with Fisher's exact test when either group has fewer than 40 respondents. Two answers from the same respondents, such as two platforms, are compared with the exact McNemar test, and two sets of 1-to-5 ratings with the Mann-Whitney U test. Where the report ranks answers to one question without saying whether the gap can be called, the ranking rests on the figures alone.

The July form accepted more than one answer on some single-answer questions, so July comparisons use only the respondents who gave one answer. Some July figures therefore differ from the July report: rigorous tracking is 46% here against 47% there, and utilization at 50% or below is 76% here against 69% there, because the July report counted all respondents operating GPUs.

The respondent mix changed between months. Technology companies fell from 35% of July's respondents to 17% of August's, and organizations above 1,000 employees from 57% to 39%. The July and August comparisons do not adjust for these shifts.