The TutorMoments dataset tests whether AI tutors can identify the exact moment a student needs help. Researchers analyzed how models balance guidance with student autonomy to avoid over-helping. This benchmark exposes a gap in current LLM pedagogical reasoning. Developers can now measure if their agents disrupt learning by providing answers too quickly.