Preference alignment reduces hallucinations in multimodal models by forcing responses to match image content. Apple researchers found that standard language alignment fails to address visual inconsistencies. This study identifies specific gaps in how MLLMs process image-text pairs. Practitioners can now use these findings to refine training loops and minimize factual errors in vision tasks.