The TutorMoments dataset tests whether AI tutors recognize the optimal moment to intervene during student learning. Researchers analyzed how models balance providing hints versus allowing productive struggle. This benchmark reveals that current LLMs often over-help too quickly. Practitioners can use these findings to refine pedagogical prompting for more effective educational agents.