Hallucinations in multimodal models often stem from responses that contradict image content. Apple Machine Learning Research analyzed how preference alignment reduces these inconsistencies. The study identifies specific gaps where standard language alignment fails to fix visual errors. These findings help developers build more reliable vision-language systems that strictly adhere to provided visual evidence.