The 3.6 billion parameter Nemotron 3.5 Lightning hits 670 tokens per second. This open-weight model matches the performance of OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. Nvidia prioritizes inference speed over raw scale. Developers can now deploy high-throughput intelligence on significantly leaner hardware footprints.