The Alibaba SAIL team integrated the Triton language to optimize GPU kernel performance. This move reduces reliance on proprietary CUDA kernels by utilizing a more flexible, Python-like syntax for hardware acceleration. Developers can now write high-performance kernels more efficiently. It is an incremental step toward better hardware abstraction in large-scale model training.