The RubricForge framework evolves judging rubrics by reflecting on small sets of ground-truth labeled trajectories. This method prevents LLM judges from mistakenly crediting fluent but unsuccessful agent paths as successes. It replaces manual rubric writing with automated, outcome-grounded induction. Practitioners can now more accurately evaluate agents when executable environment rewards are unavailable.