The RubricForge framework evolves text-based judging rubrics using reflective evolution against ground-truth labeled trajectories. This method prevents LLM judges from awarding high scores to fluent but unsuccessful agent paths. It replaces hand-written rubrics with data-grounded criteria. Practitioners can now more accurately automate agent evaluation without relying on expensive executable environment rewards.