Apple developed RayRoPE to encode image patches using predicted points along rays. This method enables SE(3)-invariant attention, solving a gap where prior absolute and relative encoding schemes failed. The approach allows transformers to better adapt to 3D scene geometry. It provides a more precise spatial framework for multi-view vision models.