Apple researchers developed Internalized Visual Thinking (IVT) to eliminate the inference overhead of visual chain-of-thought. The framework trains models to reason visually during training but output direct answers at inference. This removes the need to generate intermediate images. Practitioners gain faster proactive video reasoning without sacrificing spatial or temporal accuracy.