Six days and $3,000 in credits failed to produce a single acceptable paper from Claude Opus or GPT-5.6 Sol. Original NeurIPS authors rated all AI-generated submissions as rejects. While models handled engineering tasks, they lacked the judgment to abandon failed hypotheses. This gap suggests autonomous scientific discovery remains distant for current agentic workflows.