Our new paper shows that RoPE—the positional encoding used in most modern LLMs like Qwen, Gemma, DeepSeek—has a fundamental flaw: it entangles "what" (content) and "where" (position) information.
Our fix (PoPE) is simple but powerful. Paper:
Presenting PoPE today at #ICML2026!
We revisit RoPE through the lens of content-position entanglement, and show how polar coordinates can better decouple content from position.
Come by poster #4012 if you’re curious positional embeddings, pretraining or length generalization.
Our new paper shows that RoPE—the positional encoding used in most modern LLMs like Qwen, Gemma, DeepSeek—has a fundamental flaw: it entangles "what" (content) and "where" (position) information.
Our fix (PoPE) is simple but powerful. Paper: arxiv.org/abs/2509.10534
Interesting work! Their central result that RoPE struggles to distinguish positions and tokens in long contexts is very closely related to our PoPE paper, where we analysed RoPE entangles “what” and “where” in attention and proposed a decoupled alternative.
PoPE: arXiv 2509.10534
Excited to share our new paper: RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
LLMs often fail on inputs well within their advertised context lengths. We show that these failures are not merely engineering issues, but from intrinsic limitations of
Our new paper shows that RoPE—the positional encoding used in most modern LLMs like Qwen, Gemma, DeepSeek—has a fundamental flaw: it entangles "what" (content) and "where" (position) information.
Our fix (PoPE) is simple but powerful. Paper: arxiv.org/abs/2509.10534
Come visit our poster East Exhibit Hall A-C #3707, today (Thursday) between 4:30-7:30pm to learn about how complex-valued NNs perform perceptual grouping. #NeurIPS2024
Excited to present "Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery" at #NeurIPS2024!
TL;DR: Our model, SynCx, greatly simplifies the inductive biases and training procedures of current state-of-the-art synchrony models. Thread 👇 1/x.