A new framework uses active inference to balance the cost of clarifying questions against the risk of incorrect assumptions. The system minimizes expected free energy to decide whether to call a tool, ask a user, or proceed. This approach reduces token waste during Optimal Question Asking. Practitioners can use this to optimize agentic efficiency.