Native language training using GRPO leaves only a small performance gap compared to English reasoning. Researchers at Apple tested this across various base models and multilingual reward systems. The findings prove that verifiable rewards scale effectively beyond English. Practitioners can now optimize non-English LLMs without relying on English-centric training pipelines.