New research demonstrates that LLMs can be trained to toggle between low, medium, and high-effort reasoning modes. By adjusting compute allocation, model developers can optimize the trade-off between latency and accuracy. This allows a single model to handle simple queries quickly while reserving deep thinking for complex logic. It streamlines inference costs for practitioners.