Epoch and METR released MirrorCode to test how AI systems handle complex, week-long programming tasks. Current models fail on the hardest challenges, highlighting a gap in autonomous software engineering. This benchmark provides a concrete metric for agentic reliability. Developers can now measure exactly where long-horizon reasoning breaks down during extended coding projects.