AgentLens replaces binary pass-fail metrics with full trajectory reviews for interactive code agents. The benchmark pairs formal verification with LLM-written critiques to analyze how agents recover from errors and use tools. This approach allows developers to diagnose specific behavioral failures.
The Signal
It moves evaluation beyond simple rankings toward granular, production-ready debugging for coding agents.