Menu

Categories

Tags

Xiaomi and PKU propose MOPD distillation to end RL's 'see-saw effect'

July 2, 2026 | Source: arxiv | AI | 162 views 0 comments

The 'see-saw effect' in large model reinforcement learning has been a stubborn problem: when you train a model to excel in multiple domains like math, coding, and instruction following, the skills interfere with each other, leaving you with a jack-of-all-trades that's worse than domain-specific experts.

Now, researchers from Peking University and Xiaomi's large model team have published a paper proposing a multi-teacher online distillation framework called MOPD (Multi-Teacher On-Policy Distillation) that aims to solve this. The idea is to train separate RL 'expert' teachers for each domain (e.g., math, code) in parallel, then distill all their knowledge into a single student model at the policy level. Specifically, the student samples from its own generation trajectories and minimizes KL divergence with the corresponding teacher at the token level.

This approach avoids the exposure bias of traditional offline fine-tuning through on-policy sampling, provides dense token-level supervision that reduces variance and improves sample efficiency, and allows each domain's teacher to be developed and hyperparameter-tuned independently.

In experiments on the Qwen3-30B-A3B model, MOPD improved the normalized composite score by 5.5 points over baselines like joint RL (Mix-RL), achieving 0.937. It nearly preserved each teacher's vertical capabilities intact. The framework has already been deployed in Xiaomi's production-grade large model MiMo-V2-Flash (309B) during post-training.

The researchers also note that distillation stability depends critically on using 'homologous teachers' — the teacher and student must start from the same SFT checkpoint. If you try to distill an out-of-domain strong model (say, a 235B teacher into a 30B student) with high KL divergence, the optimization process collapses.

Leave a Reply

Your email address will not be published. Required fields are marked *