Researchers developed a method to toggle between low, medium, and high-effort reasoning modes in LLMs. This approach allows models to allocate compute based on task complexity rather than using maximum effort for every prompt. It reduces latency for simple queries. Practitioners can now optimize inference costs without sacrificing accuracy on difficult problems.