Training models to reason in native languages yields performance nearly equal to English-centric training. Apple researchers tested Group Relative Policy Optimization across diverse base models and languages to bridge this gap. The findings prove that verifiable rewards effectively scale reasoning capabilities beyond English. This enables more efficient localized model tuning.