Pinned
🧵[1/5] Qwen3.5 demonstrates the potential of hybrid LLMs. But how well do the Linear Attention layers manage their associative memory?
Previous research indicates a low effective rank, which we show:
1. Amplifies query noise
2. Poorly conditions gradients
3. Wastes memory.





