Linear probes reveal that RL-tuned models predict answer correctness more accurately than SFT counterparts. This suggests RL creates more linearly separable and structured internal representations. The researchers used mean ablation to identify a hierarchical architecture in deeper layers. These findings provide a mechanistic explanation for why reinforcement learning outperforms supervised fine-tuning in complex math tasks.