Just as its K3 model launches, Moonshot AI is already hunting for chips for the next-generation Kimi K4. The new model will be significantly larger than the 2.8 trillion-parameter K3, but there's no training or release timeline yet.
According to The Information, K3 used Nvidia's most advanced Blackwell chips, with some training actually happening in China — not entirely at the Thailand data center U.S. officials had pointed to.
More critically, K3 didn't have a single large training cluster. No Chinese cloud provider could supply enough Blackwell chips, so Moonshot AI patched together servers from at least two vendors, linking many eight-GPU machines.
With chips from different suppliers scattered across data centers, communication speed and stability became a nightmare. Moonshot's engineering team had to redesign the network to get those fragmented computing resources to train a 2.8 trillion-parameter model together.
Once trained, running K3 also requires massive compute. Moonshot primarily uses Nvidia H20 chips for inference, recommending at least 64 chips per deployment. After K3 launched, demand surged so fast that the company halted new subscriptions in under 48 hours.
Domestic chips won't fill the gap anytime soon. Large servers that can link dozens or hundreds of homegrown AI chips remain in short supply, with delivery times stretching up to six months.
Moonshot has already proven it can piece together K3 under chip constraints. For K4, there's only a direction, not a timeline. How big it can get and when training begins may depend on how many more Blackwell chips Moonshot can scrounge up.