Reinforcement Learning with Verifiable Rewards (RLVR) uses GRPO to scale model capabilities through simple trace sampling. The author argues that safety researchers often ignore these capability trends. This gap hinders the development of safety mechanisms derived from learning algorithms rather than additive loss terms. Practitioners must integrate capability research into safety frameworks.