A custom 10-room text adventure benchmark was solved for the first time by Claude Opus 5. The test requires agents to collect three specific keys to unlock a final exit. This result demonstrates improved long-term planning and state tracking. Practitioners can use such niche benchmarks to stress-test agentic reasoning beyond standard datasets.