The TutorMoments dataset tests whether AI tutors recognize the precise moment a student needs help. Researchers analyzed how models balance guidance with autonomy to avoid over-helping. This benchmark reveals that current LLMs often struggle with pedagogical timing. Developers can now measure if their agents provide hints too early or too late.