Cursor has open-sourced Mixture-of-Kittens (MoK) — a feline take on Mixture-of-Experts — to speed up MoE model training. The trick: instead of handling GPU data transfers and computation as separate steps, MoK folds them into a single kernel, the low-level program that runs directly on the GPU.
If you've ever wondered why MoE training can feel slow, here's the culprit. MoE models split into a huge number of "experts," and each forward pass only calls a subset of them. Those experts are scattered across different GPUs, so data has to shuffle between them constantly. In some cases, that data movement eats up more than half of total training time.
MoK lets the GPU move data and do math at the same time, cutting way down on idle waiting. In real training runs across 512 Nvidia GB300 GPUs, Cursor says overall throughput jumped 41%. When testing just the MoE layers in isolation, MoK was up to 2.37x faster than existing public approaches.
MoK is already running in Composer, the training stack Cursor uses across tens of thousands of GPUs. It's now available under an Apache 2.0 license. One limitation: it only supports Nvidia Blackwell GPUs right now, so it's aimed at organizations running GB200 or GB300 NVL72 clusters.