RayRoPE: Projective Ray Positional Encoding for Multi-View Attention
We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency similarity, and can be adaptive to the geometry of the underlying scene. We find that prior (absolute or relative) encoding schemes for multi-view attention do …
Read more “RayRoPE: Projective Ray Positional Encoding for Multi-View Attention”