Natural language autoencoders are the discovery that most makes me question my research intuition
You jointly train two models: one to turn activations into English, and one to turn English into activations.
The objective for the second model is to invert the first one, and the
New Anthropic research: Natural Language Autoencoders.
Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read.
Here, we train Claude to translate its activations into human-readable text.




