Linear probes reveal that RL-tuned models achieve higher accuracy in predicting answer correctness than SFT models. This suggests reinforcement learning creates more linearly separable internal representations. Mean ablation studies further show a hierarchical architecture in deeper layers. Practitioners can now better distinguish how RL improves mathematical problem-solving over standard supervision.