Researchers at Apple analyzed how preference alignment reduces hallucinations in multimodal models. The study finds that alignment prevents models from stating facts that contradict image content. This technical deep-dive clarifies how to synchronize visual and textual data. Practitioners can now better target consistency errors during the fine-tuning phase of MLLMs.