Iterative deployment of long-running models revealed new failure modes in OpenAI's safety systems. The lab identified specific risks that emerge only during extended task execution. These findings lead to updated safeguards for autonomous behavior. Practitioners can now better predict where long-horizon agents fail before full-scale release to prevent unpredictable system drift.