Iterative deployment of long-running models revealed new failure modes in OpenAI's safety systems. These long-horizon tasks create unique alignment risks that standard short-term benchmarks miss. The lab implemented improved safeguards to mitigate these specific vulnerabilities. Practitioners must now prioritize longitudinal monitoring over static evaluation to ensure stability in autonomous, multi-step agentic workflows.