Epoch and METR released MirrorCode to test AI performance on complex, week-long programming tasks. Current systems fail the hardest challenges, proving that long-horizon autonomy remains elusive. This gap highlights the distance between simple code generation and reliable software engineering. Developers should view current agentic claims with skepticism until these benchmarks improve.