The TutorMoments dataset tests whether AI tutors recognize the optimal moment to intervene during student learning. Researchers evaluated if models provide hints too early or fail to support struggling users. This benchmark forces developers to move beyond simple correctness. It prioritizes pedagogical timing over raw answer accuracy for educational LLMs.