ByteDance’s Seed team has open-sourced Cola DLM, a continuous latent diffusion language model that sidesteps the standard left-to-right token generation of large language models. Instead, it organizes high-level semantics first, then maps them to concrete text.
At its core, Cola DLM uses a Text VAE + block-causal DiT. The Text VAE maps discrete text into a continuous latent space. The block-causal DiT learns the latent prior via Flow Matching. Finally, a conditional decoder reconstructs text from the latent variables. The diffusion process operates on latent semantic representations, not directly denoising tokens.
The open-sourced version is a ~2.3B parameter model: 1.8B for the core DiT and 0.5B for the VAE. On eight benchmarks (LAMBADA, MMLU, OBQA, HellaSwag, RACE, SIQA, SQuAD, Story Cloze), the paper reports that under a unified generative evaluation protocol, it achieves scaling performance competitive with same-scale AR / LLaDA baselines and obtains the best average score.
However, this is still a research checkpoint, not a ready-to-use chat model. It has not undergone instruction tuning or RLHF, and its primary purpose is to study how continuous latent diffusion can be used for text generation. The paper also shows preliminary experiments extending to unified text-image modeling, but the open-source repository only contains the text pipeline.