Frontier Models Fail Enterprise IT Benchmark | dailyai.report
23 stories from today
Agents
93d ago
Frontier Models Fail Enterprise IT Benchmark
No frontier model scored above 50% on ITBench-AA, a new evaluation for agentic IT tasks. IBM and Artificial Analysis tested models on complex system administration and troubleshooting workflows. These results expose a critical gap between general reasoning and reliable enterprise execution.
The Signal
Practitioners should expect significant failure rates when deploying autonomous agents for technical infrastructure management.