The TutorMoments dataset tests whether AI tutors recognize the optimal moment to intervene during student learning. Researchers analyzed how models balance providing hints versus allowing productive struggle. This benchmark exposes a common failure: AI often over-helps too quickly. Educators can now quantify if a model disrupts the learning process or supports it effectively.