New research identifies how LLMs toggle between low, medium, and high-effort reasoning modes. By manipulating these levels, researchers can optimize the trade-off between compute cost and answer accuracy. This allows developers to trigger deep thinking only for complex queries. It reduces wasted inference cycles without sacrificing performance on difficult tasks.