DeepMind’s new Gemini 3.1 Flash‑Lite sets a new benchmark for speed and affordability in large‑language models. Built on the same architecture as its predecessors, it delivers faster inference with a 30% lower cost per token.
The Signal
The upgrade promises to accelerate real‑world applications, from chatbots to data‑analysis tools, while keeping power consumption minimal.