Presented by Virtana


AI infrastructure failures are hard to diagnose because a fault in one system produces symptoms in several others. When a training job slows down, the GPU dashboard might show contention, the data pipeline report a stall and the storage system start throwing I/O alerts, and each alert suggests a different fix. Infrastructure teams see every signal, and they spend hours working out where the failure started while expensive GPU capacity sits idle.

That search gets harder as enterprises build AI factories that span GPUs, storage, networks, data pipelines, and hybrid environments, because most AI factory monitoring still relies on a separate tool for each domain. And Virtana’s AI Factory Reality Check surveys of U.S. and U.K. enterprise decision makers found that 59% of U.S. enterprises and 53% of U.K. enterprises cannot automatically identify the root cause of an AI workload failure across infrastructure domains.

“Once you get to causality, you can remediate,” says Paul Appleby, president and CEO of Virtana. “If you can’t get to causality, you’re really throwing a dart at a wall to work out what happened and how to fix it.”

Scale makes AI failures harder to trace

Pressure from boards pushes enterprises to show AI progress through deployed infrastructure, so many build first and instrument later, if at all. The legacy tools that handle most of that instrumentation assume stable, bounded systems with predictable service topologies, but in an AI factory, schedulers keep moving workloads across hybrid environments, and resource consumption rises and falls in nonlinear bursts.

Threshold-based alerting needs a baseline of normal performance to measure against, and 66% of U.S. enterprises run AI infrastructure without reliable performance baselines. Only 34% of U.S. and 26% of U.K. enterprises describe AI workload performance as highly predictable, which falls to 25% at U.S. organizations with more than 50,000 employees. That leaves the companies running the largest AI factories with the least ability to anticipate what those factories will do next.

To contain premium AI hardware costs, many enterprises in both countries are rebalancing workloads across hybrid environments and consolidating systems to improve per unit efficiency, all while those systems run under load. Each of those moves alters dependencies and resource contention across the stack.

“Without system-level observability, organizations can’t determine how these changes affect outcomes, cost or reliability,” Appleby says. “They’re continuously optimizing AI systems they don’t fully understand, and introducing risk with every change.”

Tracing root cause takes a model of the entire AI factory

Root cause analysis has to establish where in a distributed system an alert originates, which in an AI factory means searching across GPUs, CPUs, memory, storage, networks, orchestration, and data pipelines at the same time.

Far more U.K. enterprises have automated detection than diagnosis. Seventy-five percent rely on automated alerting as their first response to an AI workload failure, while 47% can automatically identify the root cause across all infrastructure domains. The remainder see only a single domain, correlate signals by hand across tools, or pull multiple teams together for hours or days. In the U.S., 25% of enterprises start incident response with manual investigation across disconnected consoles.

“A storage bottleneck degrades a data pipeline, which stalls a training job, which produces GPU contention,” Appleby says. “Legacy monitoring registers three separate events in three separate domains. Each event is technically accurate, and none of them identifies what actually happened.”

Adding instrumentation to each domain leaves those three events disconnected. AI factories need a unified operational model of the full AI execution system that correlates continuous, high-fidelity telemetry from every domain in real time, including orchestration and application behavior. The model also has to maintain a live topology of how workloads, services, and infrastructure depend on each other, which lets it trace GPU contention back to the storage bottleneck that started it. That topology changes every time a scheduler moves a workload, so the model has to update continuously.

Leaders in both countries ranked unified visibility across AI and infrastructure first and AI-powered root cause analysis without manual correlation second, and those two options led in every role group and revenue band. Observability is already a multibillion-dollar market, yet adding more domain-specific monitoring doesn’t solve the cross-domain visibility and causality problems the research identifies.

AI agents need causality before they can remediate

Large AI factories will need autonomous remediation, because the throughput and complexity of those environments already exceed what manual coordination can reliably handle. Only 23% of the U.K. infrastructure and site reliability engineering (SRE) practitioners who run AI workloads day to day describe performance as highly predictable. Enterprises have to put that telemetry and live topology model in place before any agent acts on those systems, Appleby explains.

“Without that foundation, AI agents inherit the same blind spots that constrain human operators, and they amplify those failures at machine speed,” he says. “An agent that acts on incomplete system context doesn’t resolve incidents faster, it creates new ones.”

Virtana applies that approach through its Agentic Observability platform. It maintains a shared model of the full AI execution stack across on-premises, virtualized and public cloud environments. Its autonomous agents correlate GPU utilization, token demand, model behavior, and underlying infrastructure performance in real time, and they back each root cause finding with evidence.

Cost attribution and audit trails draw on the same telemetry

In both countries, 31% of enterprises call clearer ROI metrics from existing AI investments their most important prerequisite for scaling further.

“Organizations that can’t measure GPU utilization, cost per workload and efficiency at the infrastructure level can’t make the case for more investment,” Appleby says. “They can observe spending, but they can’t govern it.”

Executives and the engineers who run AI infrastructure also give different answers about whether their organizations can diagnose failures. In the U.K., 59% of executives say their organization automatically identifies root cause across all infrastructure domains, compared with 34% of the infrastructure and SRE engineers who field the alerts, a wider gap than the U.S. survey found. IT leadership, which holds final authority over AI investment in most U.K. organizations, reports the most confidence in diagnosis. That is a governance problem, because the people approving AI budgets are working from an assessment of diagnostic readiness that their own engineers don’t share.

U.K. enterprises also deploy AI under U.K. GDPR and sector oversight in financial services, healthcare, and critical national infrastructure, and those frameworks require audit trails and cost attribution. Even so, 39% of U.K. enterprises are deprioritizing security and compliance reviews as AI factory demands grow. Sovereignty will become a bigger AI issue outside the U.S., as countries and citizens push for control over their own destiny, he says.

“The system that proves an AI factory is performing is the same system that satisfies a regulator, an auditor or a board inquiry,” Appleby says. “Sovereignty without observability is a claim that can’t be evidenced.”

Enterprise AI factories are earlier than the spending suggests

The public narrative has Global 2000 companies adopting AI at massive scale, but most of the largest AI infrastructure investments come from hyperscalers and cloud platforms, Appleby says.

“It’s earlier than the level of investment would indicate,” he says. “There aren’t many examples anywhere in the world of large enterprises that have built, scaled and deployed industrial-scale AI services and are operating them efficiently.”

But an enterprise adding capacity before it establishes visibility and control will build a larger, more expensive version of the operational problems it already has.

“The gold rush is happening with the very large service providers,” he says. “For large enterprises, the opportunity still exists to get it right.”


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.