The vLLM high-throughput inference engine now supports Baidu Kunlun accelerators. This integration extends the library's hardware compatibility beyond NVIDIA and AMD GPUs. Developers can now deploy large language models on Baidu's proprietary silicon using a standardized API. It reduces vendor lock-in for enterprises operating within the Chinese cloud infrastructure ecosystem.