Linear probes reveal that RL-tuned models predict answer correctness more accurately than SFT counterparts. These models build more structured, linearly separable internal representations. Mean ablation studies confirm a hierarchical architecture where deeper layers refine reasoning. This suggests reinforcement learning fundamentally alters how models organize logic, offering a blueprint for improving mathematical accuracy in future architectures.