Hardware yearns for block sparse attention, yet it seems largely absent from open weight LLMs. DeepSeek developed NSA, and people speculated DeepSeek v4 would integrate it, yet it was never utilized.
We have a hypothesis as to why.
We found that replacing dense attention with
Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff?
We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer
This is Google’s new diffusion LLM, DiffusionGemma’s denoising canvas over time.
Diffusion LLMs can generate tokens in flexible order. But in practice, do they just become autoregressive anyway?
1/ 🧵