Pinned
NLAs look like they "work"---but compared to what? How can we know that they actually tell us anything about model internals?
In my new blogpost (and paper, to be presented at ICML!), we investigate whether activation verbalizers (like NLAs) produce faithful explanations.
New Anthropic research: Natural Language Autoencoders.
Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read.
Here, we train Claude to translate its activations into human-readable text.




