Researchers unveiled DEAF, a diagnostic benchmark with 2,700 conflict stimuli that probe whether Audio MLLMs truly process sound or merely infer from text. By varying semantic content, misleading prompts, and background noise, the framework isolates acoustic fidelity across prosody, background sounds, and speaker identity.
The Signal
The study reveals that many models still lean heavily on textual cues.