Teacher models fail 57% of trials on the τ 2-bench, wasting costly data during tool-calling distillation. Apple's new PROOF-Gen framework recovers these failures by optimizing trajectories instead of simply filtering them. This approach turns near-misses into training signals. Practitioners can now reduce frontier-model costs while improving agent reliability in complex scenarios.