Need to train a smaller AI model by learning from a bigger one? You've probably run into a wall: conventional distillation keeps both the student and teacher models in GPU memory at the same time, which eats VRAM fast. MIT and NVIDIA have a fix called Lightning OPD (Offline On-Policy Distillation).
https://twitter.com/hancai_hm/status/2054071440761409987
Instead of keeping the teacher model running in real-time during training, Lightning OPD precomputes the teacher's log-probabilities offline. That frees up all the GPU memory for the student model alone, boosting training speed by 4x. In tests on a single node with 8 H100s, the team successfully distilled the reasoning capabilities of a 30B-parameter model (Qwen3-30B-A3B-Thinking) into a base 30B model (Qwen3-30B-A3B-Base), achieving a score of 71.0 on the AIME 2024 benchmark. Standard OPD would have simply run out of memory trying to fit both 30B models.
The framework also showed efficiency in a 32B-to-8B distillation run, hitting a score of 69.9 in just 30 GPU hours.
The researchers note a hidden prerequisite for offline distillation: teacher consistency. The student model must use the same teacher model during both supervised fine-tuning (SFT) and subsequent distillation stages. Violating this principle causes gradient misalignment and ultimately degrades performance.