The RubricForge system evolves judging rubrics by reflecting on small sets of ground-truth labeled trajectories. This method prevents LLM judges from mistakenly crediting fluent but unsuccessful agent paths as successes. It replaces manual rubric writing and weight fine-tuning. Practitioners can now automate high-fidelity agent evaluation without expensive executable environment rewards.