Researchers unveiled a novel expert‑prefetching strategy that predicts which MoE components will be needed next, allowing memory transfers to overlap with computation. By reducing CPU‑GPU bottlenecks, the technique boosts decoding speed for large language models while keeping memory usage low.
The Signal
This advance supports broader deployment of scalable AI systems across cloud, edge, and research settings worldwide.