Three evaluation axes—composition, grounding, and instruction-following—drive a new reward framework from Apple. The system generates query-specific rubrics based on retrieved evidence to provide fine-grained supervision during post-training. This replaces holistic scalar objectives with decomposed quality dimensions. Practitioners can now better align open-domain question answering models with specific, verifiable knowledge constraints.