Menu

Categories

Tags

Nous open-sources Lighthouse Attention, speeding up 512K context by 17x on a single B200

May 18, 2026 | Source: nousresearch | AI, NVIDIA | 154 views 0 comments

Nous Research has open-sourced Lighthouse Attention, a long-context pretraining mechanism. On a single B200 GPU processing 512K-token sequences, the method delivers roughly 17x faster computation than traditional attention, and at 98K tokens it achieves 1.4 to 1.7x end-to-end training speedups.

Standard attention computes pairwise relationships between every token, so compute costs explode quadratically as sequence length grows. Lighthouse Attention takes a "coarse filtering then fine computation" approach: it first quickly scans compressed summaries of the text at different levels, scores and picks out core segments, concatenates them into a short sequence, and then feeds that directly to the efficient FlashAttention kernel. Because the filtering logic is completely decoupled from the kernel, developers don't have to write low-level code or add extra training objectives.

Past acceleration schemes with a similar philosophy often came with side effects — models that learned to skim reading often lost their ability to read word-by-word. To avoid that trap, the research team lets the model run in accelerated mode for most of training, and only briefly switches back to full traditional attention at the very end as a kind of adaptation. In tests on a 530-million-parameter model trained on 50 billion tokens, the resulting model not only cut training time significantly but also matched or outperformed a baseline trained entirely with standard attention.

Tags: #Github

Leave a Reply

Your email address will not be published. Required fields are marked *