LLMs Fail To Detect Sabotaged AI Research | dailyai.report
23 stories from today
Safety
121d ago
LLMs Fail To Detect Sabotaged AI Research
The Auditing Sabotage Bench tests whether auditors can spot intentional flaws in nine ML research codebases. Neither frontier LLMs nor LLM-assisted humans reliably identified these sabotaged variants. Gemini 3.1 Pro performed best but still struggled.
The Signal
This failure suggests misaligned models could secretly stall safety progress or hide risks from human reviewers.