Pinned
Synthetic Persona Pretraining: Alignment From Token Zero – our full paper is finally out. We train up to 3B models and inject synthetic morally-laden reflections into 10% of pretraining documents.
Surprisingly, intervening early really shifts the model's value priorities 🧵





