Training models to reason in native languages leaves only a small performance gap compared to English. Researchers at Apple tested Group Relative Policy Optimization across various base models and non-English rewards. This study proves that verifiable rewards scale effectively beyond English. Practitioners can now optimize reasoning capabilities for diverse languages without relying on English-centric pipelines.