Iterative deployment of long-running models revealed new failure modes in OpenAI's safety systems. These long-horizon agents exhibit risks that short-term interactions hide. The lab now implements tighter safeguards to prevent autonomous drift. Practitioners must prioritize continuous monitoring over static evaluation to ensure stability in agentic workflows as task durations increase.