Neural networks do math by rotating shapes.
We found a shape-rotating calculator hidden inside an LLM โ and itโs used for more than just math! (1/6)
Submit your work! The 2nd Workshop on ๐๐๐ญ๐ข๐จ๐ง๐๐๐ฅ๐ ๐๐ง๐ญ๐๐ซ๐ฉ๐ซ๐๐ญ๐๐๐ข๐ฅ๐ข๐ญ๐ฒ will be held at COLM 2026 in San Francisco!
Submission Deadline: June 21, 2026
@ActInterp
The most popular way to interpret AI is missing the bigger picture.
Models think in curved shapes. But sparse autoencoders (SAEs) work with straight lines.
Can they still capture modelsโ curved neural geometry? Yes, but not how you might think! (1/7)
The demo below blows my mind.
1) This is not a programmed calculator. This is a โcalculatorโ that emerged inside a specific layer of a language model.
2) Itโs so cool that you can extract it from the model and present it this clearly.
The same calculator handles a wide range of tasks, including:
- arithmetic (โ7+9โ)
- weekdays (โnine days after Fridayโ)
- months (โsix months after Augustโ)
Llama built this mechanism from scratch in training, and uses it with striking elegance and flexibility. (4/6)
@sheridan_feucht explaining in depth what Iโve been doing for the past 3 months ๐งฎ๐ฆ
I know Iโm not objective, but I think we have beautiful results in this paper!
Neural networks have beautiful feature geometry, but do they have mechanisms that actually interface with those structures?
At @GoodfireAI this spring, we discovered one: a re-usable addition mechanism that reads/writes to Fourier features from prior work. ๐งต