SPEED-Bench is a new unified benchmark that evaluates speculative decoding across diverse models and tasks. Developed by Hugging Face, it aggregates metrics for latency, accuracy, and resource usage, enabling researchers to compare techniques like top‑k sampling and beam search.
The Signal
The benchmark promotes reproducibility and helps teams prioritize efficient inference pipelines.