Native language reasoning training leaves only a small performance gap compared to English training. Apple researchers tested Group Relative Policy Optimization across various base models and languages to verify these results. The study proves that verifiable rewards scale effectively beyond English. Practitioners can now optimize multilingual reasoning without relying on English-centric pipelines.