Menu

Categories

Tags

xAI Has 500,000 GPUs but Uses Only 11% of Them. Rivals Say That's 'Ridiculously

April 30, 2026 | Source: theinformation | AI, xAI | 997 views 0 comments

Elon Musk's xAI has around 500,000 Nvidia GPUs — one of the largest clusters for any AI developer by public counts. But according to an internal memo, the company's MFU (Model Flops Utilization, a measure of how much of a chip's theoretical peak performance is actually being used) has hovered around 11% in recent weeks. A researcher at a competing lab says most companies struggle to crack 40%, but 11% is "ridiculously low."

Low utilization is a known problem across the industry. AI training is stop-and-go: GPUs run flat out while training, then sit idle while researchers analyze results and decide what to do next. There are hardware bottlenecks too — high-bandwidth memory (HBM) can't keep pace with compute cores, and any weak link in the network connecting thousands of GPUs can slow the whole cluster down. The hunger for AI training chips is insatiable — Amazon's Trainium2 is already sold out — but actually using them efficiently is another story.

There's also a "gaming the numbers" phenomenon. A researcher at a large lab says colleagues will rerun training experiments over and over to boost utilization figures — partly to avoid criticism from management, and partly to prevent idle GPUs from being reassigned to other teams.

Leave a Reply

Your email address will not be published. Required fields are marked *