Three evaluation axes—composition, grounding, and instruction-following—drive a new reward framework from Apple. The system generates query-specific rubrics grounded in retrieved evidence to replace holistic scalar objectives. This decomposition provides finer supervision during post-training. Practitioners can now optimize for multiple quality dimensions simultaneously rather than relying on a single, vague reward signal.