Apple researchers analyzed how preference alignment reduces hallucinations in multimodal models. These systems often produce text inconsistent with image content. The study focuses on forcing responses to adhere strictly to visual data. This technical deep-dive provides a framework for developers to minimize factual errors in image-to-text tasks.