Researchers unveiled a test‑time quantization method that compresses large language models during inference, eliminating the need for costly retraining. By calibrating on the fly, the technique adapts to any prompt, reducing latency across diverse tasks.
The Signal
Early experiments show speed gains surpassing existing approaches, promising broader deployment of powerful AI systems worldwide.