Hallucinations in multimodal models often stem from responses that contradict visual input. Apple researchers analyzed how preference alignment reduces these inconsistencies in image understanding tasks. The study identifies specific failure modes where models ignore visual evidence for linguistic patterns. This provides a technical framework for developers to improve visual grounding in MLLMs.