Recent third-party cybersecurity evaluations revealed gaps in how OpenAI tests its models. The company is now implementing stricter safeguards to prevent model misuse during red-teaming exercises. These updates aim to harden system boundaries. Practitioners should expect more rigorous constraints on how models interact with external security tools during future evaluations.