Menu

Categories

Tags

MIT and NVIDIA's Lightning OPD speeds up AI training 4x without VRAM blowouts

May 12, 2026 | alex | AI, NVIDIA | 152 views 0 comments

Need to train a smaller AI model by learning from a bigger one? You've probably run into a wall: conventional distillation keeps both the student and teacher models in GPU memory at the same time, which eats VRAM fast. MIT and NVIDIA have a fix called Lightning OPD (Offline On-Policy Distillation).

Instead of keeping the teacher model running in real-time during training, Lightning OPD precomputes the teacher's log-probabilities offline. That frees up all the GPU memory for the student model alone, boosting training speed by 4x. In tests on a single node with 8 H100s, the team successfully distilled the reasoning capabilities of a 30B-parameter model (Qwen3-30B-A3B-Thinking) into a base 30B model (Qwen3-30B-A3B-Base), achieving a score of 71.0 on the AIME 2024 benchmark. Standard OPD would have simply run out of memory trying to fit both 30B models.

The framework also showed efficiency in a 32B-to-8B distillation run, hitting a score of 69.9 in just 30 GPU hours.

The researchers note a hidden prerequisite for offline distillation: teacher consistency. The student model must use the same teacher model during both supervised fine-tuning (SFT) and subsequent distillation stages. Violating this principle causes gradient misalignment and ultimately degrades performance.

Tags: #Qwen

Leave a Reply

Your email address will not be published. Required fields are marked *