Iterative deployment of long-running models revealed new failure modes in OpenAI's latest systems. These long-horizon tasks introduce risks that short-term interactions typically mask. The lab implemented improved safeguards to mitigate these emergent behaviors. Practitioners must now prioritize continuous monitoring over static safety checks to prevent autonomous drift during extended execution windows.