Unsolved mathematical conjectures now serve as the ultimate benchmark for LLM reasoning. Current models struggle with proofs requiring deep logical leaps rather than pattern matching. This gap highlights a ceiling in synthetic data scaling. Researchers must now find ways to instill formal verification to move beyond probabilistic guessing in high-stakes mathematics.