Apple researchers developed Internalized Visual Thinking (IVT) to eliminate the inference overhead of generating intermediate reasoning images. The framework trains models to simulate visual foresight internally while outputting only text. This removes the latency bottleneck of Visual CoT. Practitioners can now deploy proactive video reasoning without the heavy computational cost of image generation.