Apple researchers developed RayRoPE to encode image patches using predicted points along rays. This method enables SE(3)-invariant attention and adapts to scene geometry better than absolute or relative schemes. It solves a specific gap in how multi-view transformers handle spatial positioning. Practitioners can now achieve more precise geometric alignment in 3D vision tasks.