Apple researchers analyzed how preference alignment reduces hallucinations in multimodal models. These models often produce responses that contradict visual evidence. The study focuses on forcing outputs to align more strictly with image content. This research provides a technical framework for developers to minimize factual errors in vision-language tasks.