The RayRoPE encoding scheme represents image patches using predicted points along rays rather than simple directions. This approach enables SE(3)-invariant attention and adapts to scene geometry more effectively than absolute or relative encodings. Apple researchers designed the system to uniquely encode patches across multiple posed input images for better spatial reasoning.