Enterprises Bought $401B in GPUs — 95% Sit Idle

Rows of GPU server racks mostly idle in a dark data center
KEY TAKEAWAYS
  • Cast AI’s third annual audit of tens of thousands of Kubernetes clusters across AWS, Azure, and GCP found average enterprise GPU utilization stuck at just 5%.
  • Gartner estimates AI infrastructure will add $401 billion in new spending this year — most of it parked behind idle silicon.
  • CPU utilization fell from 10% to 8%, memory dropped from 23% to 20%, and CPU overprovisioning jumped from 40% to 69% year over year.
  • Top operators hit 49% on H200s and 30% on H100s — the gap is not hardware, it is automated rightsizing and workload scheduling.

For every dollar enterprises spend on a GPU today, ninety-five cents pays for silicon that sits idle. That is not a rhetorical flourish — it is the number Cast AI pulled from production clusters across the three big hyperclouds. Pair it with Gartner’s $401 billion AI infrastructure forecast for the year and the picture is clear: the most-hyped balance sheet item of 2026 is mostly empty capacity, and the bill is now arriving at the CFO’s desk.

The 5% Problem

A Number That Refuses to Improve

Cast AI co-founder Laurent Gil put it bluntly: “This is the third year we’ve published this report. The numbers are worse.” Average GPU utilization across the audited fleet sits at 5%, with organizations reserving roughly twenty times the GPU capacity their workloads actually consume at any given moment. CPU utilization dropped from 10% to 8%. Memory dipped from 23% to 20%. Overprovisioning — the gap between what is allocated and what is used — moved the wrong way on every axis the report measured.

Idle CPUs Are Pennies. Idle GPUs Are Dollars.

An idle CPU core leaks cents per hour. An idle H100 or H200 leaks dollars per hour, and there are millions of them sitting in production fleets. For the first time since EC2 launched in 2006, GPU prices are also moving up rather than down: AWS raised H200 Capacity Block prices by 15% in January 2026, breaking a two-decade precedent. Cost per inference and total cost of ownership rose from 34% to 41% as a top enterprise priority in a single quarter, turning utilization into a finance issue as much as an engineering one.

Business Insight — If your AI P&L looks healthy on paper, ask one question: what is the utilization rate of the GPUs we are paying for? A 5% answer means 95 cents of every silicon dollar is funding the hyperscaler, not your model. That is a procurement failure, not a roadmap question.


Why It Is Getting Worse, Not Better

The Self-Reinforcing Hoarding Loop

Most enterprises locked GPU capacity into three- to five-year depreciation cycles during the 2024–2025 scramble. Hyperscalers themselves run on five-year cycles. The result: capacity bought at the peak of the panic is now a fixed cost regardless of usage, and procurement teams refuse to release it because reorder lead times remain long. Hoarding feels safer than running out. The hoarding then tightens supply, which keeps prices elevated, which deepens the impulse to hoard. The loop has been running for two full report cycles and is getting tighter.

Outsourcing Intent Is Spiking

VentureBeat’s Q1 2026 enterprise infrastructure tracker shows the intention to evaluate inference outsourcing and managed LLM providers jumped from 13.2% to 23.1% in a single quarter. Workloads are migrating to specialized AI clouds — that share rose from 30.2% to 35.9% over the same window. The signal is unambiguous: buyers are quietly preparing to walk away from the on-prem GPUs they paid for, because the unit economics on managed inference now beat their own underused fleets.

Business Insight — The next 18 months will separate operators who treat GPU capacity as a continuously optimized portfolio from those still treating it as a one-time purchase. The former will compound margin. The latter will keep writing off depreciation on hardware nobody is using.


The Automation Gap

From 5% to 49% — What the Winners Actually Do

The Cast AI dataset is not uniform. One organization in the audit hit 49% utilization on H200s and 30% on H100s — roughly ten times the average. The difference was not better hardware, deeper pockets, or a more favorable cloud contract. It was operational: automated rightsizing, GPU sharing and time slicing, and Spot management running as continuous processes rather than annual procurement reviews. The tools to close the gap already exist. The discipline to deploy them does not.

Three Things a CEO Should Audit This Week

First, pull the average GPU utilization across every cluster you own. If it is under 20%, you have a finance problem, not a capacity problem. Second, check whether overprovisioning is measured at all — most organizations cannot report it because no one owns the metric. Third, compare your blended cost per inference against managed alternatives. The buyers who moved the line from 13% to 23% in one quarter did this math first and let the numbers decide.

Business Insight — The companies that hit 49% utilization are not smarter. They simply made resource efficiency an automated, continuous process owned by a real team. Everyone else is paying a tax on inertia, and the tax compounds every quarter the H200 prices stay up.


Related

Sources

  1. VentureBeat — 5% GPU utilization: The $401 billion AI infrastructure problem enterprises can’t keep ignoring
  2. TechRadar Pro — ‘5% utilization is a math fail’: Millions of GPUs worth billions are mostly sitting idle

AI Biz Insider · AI Business EN · aibizinsider.com


AI Biz Insider에서 더 알아보기

구독을 신청하면 최신 게시물을 이메일로 받아볼 수 있습니다.

코멘트

댓글 남기기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기