BeSafe‑Bench, a new benchmark, exposes behavioral safety risks of multimodal agents operating in real‑world web, mobile, and embodied environments. By simulating functional tasks, it reveals gaps that could affect global deployments of autonomous systems.
The Signal
The study urges developers, regulators, and researchers worldwide to adopt rigorous safety testing, ensuring reliable and trustworthy AI across industries.