This is exactly why Spiral is focusing on AI data infrastructure. Training on the right data is a surprisingly outsized lever
Pretraining progress seems to be coming mostly from data improvements.
@who_is_jerbear and I pretrained combinations of year-representative open model recipes and data corpuses across 2019 to 2025 at various small scales.
Data improvements contributed 3.24x as many compute



