New Benchmark Exposes Agent Safety Risks
By exposing hidden risks across four domains, the benchmark equips developers and regulators worldwide to design safer autonomous agents, fostering responsible AI deployment and reducing unintended harm in digital and physical environments.