Menu

Categories

Tags

Nous Research’s TST Speeds Up Pretraining 2–3x, Then a Similarity Dispute Erupts

May 14, 2026 | Source: huggingface | AI | 114 views 0 comments

Nous Research has released a new pretraining method called Token Stacking Training (TST). The technique claims to cut pretraining time by two to three times for the same compute budget by packing and compressing adjacent tokens early in training.

TST works in two phases. During the first 20 to 40 percent of training, the model doesn't process tokens one by one. Instead, it averages groups of neighboring tokens and feeds them in as a single input. At the output side, it predicts which tokens are in the next pack — without caring about their order. After that phase, the model falls back to standard next-token prediction. Because the underlying architecture isn't changed, the final model behaves identically to a conventionally trained one during inference. Nous says it has validated the method on Mixture-of-Experts models up to 10 billion parameters.

The core idea is trading data for compute: burning through training data faster to reduce wall-clock time. That could become a liability if high-quality text corpora start running low.

Within hours of the paper going live, readers pointed out that TST’s mechanism looks nearly identical to a 2024 preprint called Beyond Next Token Prediction. The TST team acknowledged the resemblance on Hugging Face, calling it an "unfortunate case of convergent research," and promised to update the paper with a citation.

Leave a Reply

Your email address will not be published. Required fields are marked *