Current benchmarks for recursive self-improvement focus on narrow, easily verifiable tasks. These metrics mislead observers into believing AI agents can automate discovery. True research remains open-ended and lacks immediate verification. This gap suggests that forecasts of explosive, automated progress are premature. Practitioners should distrust benchmarks that ignore the ambiguity of scientific discovery.