The vLLM high-throughput inference engine now supports Baidu Kunlun accelerators. This integration expands the library's hardware compatibility beyond Nvidia and AMD ecosystems. Developers can now deploy large language models on Baidu's proprietary silicon. It reduces hardware lock-in for enterprises operating within the Chinese cloud infrastructure market.