Epoch and METR released MirrorCode to test if AI systems can execute programming tasks spanning a full week. Current models fail the most difficult challenges. This gap confirms that autonomous agents still struggle with long-term planning and consistency. Developers can now quantify exactly where agentic workflows break during complex, multi-day software engineering projects.