Finding a good training-data mixture usually means training thousands of small proxy models to search for one. Michael Hu's team got it down to one run per dataset: train a LoRA adapter on each dataset, then merge the adapters, and the merge approximates what training on that
DatologyAI builds tools to automatically select and optimize the best data on which to train AI models, leading to better, smaller models which train faster.
- An AI agent broke into Hugging Face's systems on its own. What let Hugging Face catch it fast? A capable open model to defend with. @arimorcos on why powerful open models are becoming part of the defense toolkit, not just a risk, on the latest @jacobeffron roundup.
- This week's Summer of Data talk is live. Thanks to Pranjal Aggarwal (@PranjalAggarw16, Carnegie Mellon) for coming by to talk about Gym Anything - using AI agents to automatically build the RL environments that computer-use agents train on, turning a data bottleneck into 12,000+
- This week's Summer of Data talk is live. Thanks to Adhiraj Ghosh (@adhiraj_ghosh98, University of Tübingen) for coming by to talk about task-adaptive data curation, curating training batches for concept diversity, and the 3.33x compute multiplier it produced in pretraining. Full
- This week's Summer of Data talk is live. Thanks to Furong Huang @furongh (Associate Professor at the University of Maryland) for stopping by to talk about building self-improving foundation models, models that can detect, explain, and recover from their own failures, and why

