
Apple researchers ran a reinforcement learning experiment across nine models and 11 languages to answer a deceptively simple question: if a model trains on only one language, can it reuse that problem-solving ability on questions in other languages?
The answer, they found, is yes — and the effect is surprisingly pronounced. Train a model on a single language, and it gets better at many other languages too.
On French tests, for instance, training directly on French boosted scores by an average of 25.6 percentage points. But training only on Spanish — no French at all — still improved French scores by 24.6 points. That's a gap of just one percentage point.
The implication: models aren't just memorizing questions in one language. They're picking up reasoning strategies that transfer across languages. So if you want to make a model's Chinese reasoning sharper, you may not need to rebuild all your reinforcement learning data in Chinese.
But the training language isn't arbitrary either. Specific languages can trigger specific capability regressions. Run the same reinforcement learning setup with a different training language, and some models actually get noticeably worse on other languages and tasks. The most dramatic example: Qwen3-4B, after training on Swahili, dropped 19.2 percentage points on an English test it hadn't seen during training — while multilingual training improved that same test by 4.5 points.
The paper focuses on reasoning tasks where answers can be automatically graded — math, logic, graphs, geometry. Swap the language, and the underlying solution usually doesn't change. But reading comprehension, metaphor, semantic judgment, and cultural knowledge — tasks that genuinely depend on language itself — are untested here.