A new framework from Apple replaces scalar reward signals with query-specific rubrics grounded in retrieved evidence. This method decomposes answer quality into composition, grounding, and instruction-following dimensions. It provides finer supervision during post-training than holistic objectives. Practitioners can now optimize LLMs for factual precision without sacrificing response structure or adherence to complex prompts.