Pinned
checkout our new paper about the superficial alignment hypothesis :) we use algorithmic information theory to formalise this hypothesis, we unify prior work on the topic, and show how post-training affects it!
follow @tvergarabrowne for more great work like this in the future!
first paper of the phd 🥳
the Superficial Alignment Hypothesis (SAH) argues that pre-training adds most of the knowledge to a model, and post-training merely surfaces it.
however, this hypothesis has lacked a precise definition. we fix this.




