The vLLM high-throughput inference engine now supports Baidu Kunlun accelerators. This integration expands the library's hardware compatibility beyond NVIDIA and AMD GPUs. Developers can now deploy large language models on Chinese domestic silicon. It reduces vendor lock-in for enterprises operating within the Baidu cloud ecosystem and optimizes local inference performance.