Skip to main content

PoPE-pytorch

Efficient implementation (and explorations) into polar coordinate positional embedding (PoPE) - from Gopalakrishnan et al. under Schmidhuber

Install

$ pip install PoPE-pytorch

Usage

import torch
from PoPE_pytorch import PoPE

# define pope

pope = PoPE(64, heads = 8)

# pass in sequence length

pos_emb = pope(1024)

# queries and keys in attention

q = torch.randn(1, 8, 1024, 64)
k = torch.randn(1, 8, 1024, 64)

# training

rotated_q, rotated_k = pope.apply_pope_to_qk(pos_emb, q, k)

# inference

rotated_q, rotated_k = pope.apply_pope_to_qk(pos_emb, q[..., -1:, :], k)

Axial PoPE

For images, video, etc. where multiple dimensions are needed, you can use AxialPoPE. The feature dimension will be split across these axial dimensions.

You can either pass in the positions manually, or just pass the dimensions as a tuple, in which case the grid positions will be automatically generated.

import torch
from PoPE_pytorch import AxialPoPE

# axial pope for images (e.g. 16x16)
# split 64 dim into 32 (x) and 32 (y)

pope = AxialPoPE(
    dim = 64,
    heads = 8,
    axial_dims = (32, 32)
)

pos_emb = pope((16, 16)) # (256, 64) frequencies

# for video (e.g. 8 frames, 16x16 frames)
# split 96 dim into 32 (t), 32 (x), 32 (y)

pope_video = AxialPoPE(
    dim = 96,
    heads = 8,
    axial_dims = (32, 32, 32)
)

pos_emb_video = pope_video((8, 16, 16)) # (2048, 96) frequencies

# queries and keys
# then apply to q, k as usual

q = torch.randn(1, 8, 2048, 96)
k = torch.randn(1, 8, 2048, 96)

rotated_q, rotated_k = AxialPoPE.apply_pope_to_qk(pos_emb_video, q, k)

Fused Attention Similarity

import torch
from PoPE_pytorch import PoPE, compute_attn_similarity

# define pope

pope = PoPE(dim = 64, heads = 8).cuda()

# get rotations

pos_emb = pope(1024)

# queries and keys

q = torch.randn(1, 8, 1024, 64).cuda()
k = torch.randn(1, 8, 1024, 64).cuda()

# fused attention similarity, avoiding expanding 64 to 128

sim = compute_attn_similarity(q, k, pos_emb) # (1, 8, 1024, 1024)

attn = sim.softmax(dim = -1) # the usual in attention..

Fused Flash Attention

import torch
from PoPE_pytorch import PoPE, flash_attn_with_pope

# pope

pope = PoPE(dim = 32, heads = 8).cuda()

# queries, keys, values for attention

q = torch.randn(2, 8, 1024, 64).cuda()
k = torch.randn(2, 8, 1024, 64).cuda()
v = torch.randn(2, 8, 1024, 64).cuda()

pos_emb = pope(1024)

mask = torch.ones((2, 1024)).bool().cuda()

out = flash_attn_with_pope(q, k, v, pos_emb = pos_emb, causal = True, mask = mask)

assert out.shape == (2, 8, 1024, 64)

PoPE with Mixed Rotated and Unrotated Tokens

For architectures like Vision Transformers (ViT) with multiple CLS or register tokens, you may want to append unrotated tokens to your image patches. You can pass explicit position indices to specify which tokens undergo PoPE rotations. When unrotated tokens interact with rotated tokens (or other unrotated tokens), they do so without any relative positional bias.

import torch
from PoPE_pytorch import AxialPoPE, flash_attn_with_pope

# 16x16 image patches + 4 cls / register tokens

num_patches = 256
num_register_tokens = 4
seq_len = num_patches + num_register_tokens

pope = AxialPoPE(dim = 64, heads = 8, axial_dims = (32, 32)).cuda()

# generate positions for the 16x16 grid

pos_emb = pope((16, 16))

# apply positions to first 256 tokens, leave 4 register tokens unrotated

pos_indices = torch.arange(num_patches, device = 'cuda')

q = torch.randn(1, 8, seq_len, 64).cuda()
k = torch.randn(1, 8, seq_len, 64).cuda()
v = torch.randn(1, 8, seq_len, 64).cuda()

# pass indices to handle unrotated tokens

out = flash_attn_with_pope(
    q, k, v,
    pos_emb = pos_emb,
    pope_pos_emb_indices = pos_indices
)

Citations

@misc{gopalakrishnan2025decouplingwhatwherepolar,
    title   = {Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings},
    author  = {Anand Gopalakrishnan and Robert Csordás and Jürgen Schmidhuber and Michael C. Mozer},
    year    = {2025},
    eprint  = {2509.10534},
    archivePrefix = {arXiv},
    primaryClass = {cs.LG},
    url     = {https://arxiv.org/abs/2509.10534},
}

Metadata

Release files for PoPE-pytorch 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for PoPE-pytorch 0.2.1
File Size Uploaded
pope_pytorch-0.2.1.tar.gz 17.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for PoPE-pytorch 0.2.1
File Interpreter ABI Platform
pope_pytorch-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 36.7 kB

Release files / pope_pytorch-0.2.1.tar.gz

Download URL pope_pytorch-0.2.1.tar.gz
Size 17.3 kB
Tags Source
SHA-256 checksum
How to use checksums
5655aa25aa72523b49ed03f9c5ce26e0faa082017abfa641529445edbd26f823
BLAKE2b-256 checksum
How to use checksums
9d302a8f6b7e94fee71cf135052d8ab5588d2fd67a8f744d16b4324f282c836c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.17

Release files / pope_pytorch-0.2.1-py3-none-any.whl

Download URL pope_pytorch-0.2.1-py3-none-any.whl
Size 19.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b9dc493c190aa80d1ee9a4ff7343d9f9b5cbdfaace3bbf3c8b10b22d2c5c9e1f
BLAKE2b-256 checksum
How to use checksums
2672e61cf723799a3d02a6cda894f0f9f50069691d344413fd37168a3aaee3d0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.8.17

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page