Why what you attend to can't be static:
sethmorton.com/blog/what_you_…
Transformers can't change what they attend to after training. Backprop is too global and too destructive for continual learning. The brain doesn't work this way. 1/2
The Geometry of Surprise: sethmorton.com/blog/the_geome…
Reward models are trained on a snapshot of the world. For continual learning that's a dead end - the model has to learn from its own surprise.
The question is what shape that surprise takes. 1/2
self-supervised loss is not talked about enough - incredibly cool paper. the problem is raw compute cost & memory cost -- this alludes to @carnot_cyclist 's blog, one of tenants of continual learning is efficiency. once again, thermodynamic computation is needed