Linear probes reveal that Reinforcement Learning models create more structured, linearly separable hidden states than supervised fine-tuned versions. This structural advantage allows deeper layers to better predict answer correctness during mathematical tasks. The findings clarify why RL outperforms SFT. Practitioners can now use these representational metrics to evaluate reasoning quality before final inference.