Alibaba's SAIL group now utilizes the Triton language to optimize GPU kernel performance. This move reduces reliance on proprietary CUDA code by enabling higher-level programming for deep learning operations. Developers can now write more portable and efficient kernels. This is an incremental shift toward open-standard hardware abstraction in large-scale AI infrastructure.