The RubricForge framework evolves judging rubrics by reflecting on small sets of ground-truth labeled trajectories. This method prevents LLM judges from mistakenly crediting fluent but unsuccessful agent paths. By grounding the text in true outcomes, it replaces manual rubric writing. Practitioners can now automate high-fidelity agent evaluation without expensive environment rewards.