Nvidia has open-sourced a discrete text diffusion architecture called Nemotron-Labs-TwoTower, aimed at breaking the bottleneck of large language models generating text one token at a time. Previous text diffusion models tried to achieve parallelism by forcing a single network to handle both unidirectional context understanding and bidirectional error correction simultaneously, which significantly degraded cognitive performance.
https://twitter.com/NVIDIAAI/status/2072394812301480067
TwoTower uses a decoupled dual-tower design: one tower is a completely frozen pretrained autoregressive LLM, serving as a "read-only context tower" that preserves full reasoning and common sense; the other is a separately trained "denoising writer tower" that reads context information at the layer level via cross-attention.
The writer tower employs a "confidence-based demasking" mechanism: when predicting a block, it first writes high-confidence tokens, then progressively fills in the remaining blanks, enabling easy-to-hard parallel generation. On a 30B-class hybrid architecture (Mamba-Transformer MoE), this design retains 98.7% of quality after adapting with only 2.1 trillion tokens (1/12 of the baseline model's pretraining data) and achieves a 2.42x real-world speedup without additional cache overhead. However, because both towers must remain in memory, static VRAM usage increases, and there is slight accuracy degradation in highly complex code and math reasoning tasks.