Researchers developed a method to toggle between low, medium, and high-effort reasoning modes in LLMs. This approach allows models to allocate compute based on task complexity rather than using maximum effort for every query. It optimizes inference costs. Practitioners can now balance response latency against accuracy by selecting specific reasoning depths.