Menu

Categories

Tags

Kaiming He's ELF: The language diffusion model finally works

May 13, 2026 | Source: arxiv | AI | 136 views 0 comments

MIT Kaiming He's team has released ELF (Embedded Language Flows), a language diffusion model that ditches the GPT-style autoregressive "predict next token" approach. Instead, it generates text entirely in a continuous embedding space, only converting to discrete tokens at the very last step.

Diffusion models are old hat for images, but they've always felt clumsy with language. That's because pixels are smooth gradients, but words are discrete Lego bricks. Previous continuous diffusion text models either repeatedly injected token-level supervision into the generation trajectory or required a separate decoder. ELF's approach is cleaner: it spends most of its time denoising in a continuous vector space, and only flips to discrete tokens at the very end, using a shared-weight network to make the jump.

The results are striking. On OpenWebText unconditional generation, the 105M-parameter ELF-B model achieves a Gen. PPL of about 24.1 with 32 sampling steps, beating multiple discrete and continuous diffusion language model baselines. More importantly, ELF-B was trained on only about 45 billion tokens — roughly an order of magnitude less than the 500 billion-plus tokens used by comparable methods. This suggests that the continuous diffusion path isn't fundamentally blocked by language discreteness; the real issues likely lie in modeling interfaces and sampling design.

Leave a Reply

Your email address will not be published. Required fields are marked *