Apple study finds GRPO training works in non-English languages

Apple ML Research's large-scale study tests GRPO-based RLVR across many base models and languages, finding native-language reasoning training leaves only a small gap to English. It also shows strong crosslingual transfer, but warns that some languages cause severe out-of-domain regressions, requiring broad evaluation.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- Z.ai CEO Jie Tang: GLM 5.3 gains come from RL, not parameter count
- New tool adds 14 skills to Claude Code and Cursor for Markdown diagrams
- Tool turns Claude into a team of AI employees on your Mac
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills