RayRoPE represents image patch positions using rays and predicted points rather than simple directions. This approach enables SE(3)-invariant attention and adapts to specific scene geometries. Apple researchers developed the method to fix gaps in existing absolute and relative encoding schemes. It provides a more precise spatial framework for transformers processing multiple posed input images.