The vLLM high-throughput inference engine now supports Baidu Kunlun accelerators. This integration expands the library's hardware compatibility beyond NVIDIA and AMD GPUs. Developers can now deploy large language models on Kunlun silicon using vLLM's PagedAttention mechanism. It is an incremental update that broadens the open-source ecosystem for Chinese hardware.