The newly released BeSafe‑Bench benchmark exposes behavioral safety risks of LMMs operating in real‑world settings. By testing agents across web, mobile, and embodied visual‑language domains, it reveals gaps that previous low‑fidelity evaluations missed.
The Signal
The framework equips researchers and industry worldwide to assess, improve, and standardize safety for autonomous decision‑making systems.