The TutorMoments dataset tests whether AI tutors recognize the precise moment a student needs help. Researchers analyzed how models balance guidance with student autonomy to avoid over-helping. This benchmark exposes a gap in current LLM pedagogical reasoning. Developers can now measure if their agents hinder or help the actual learning process.