Like (nearly) everyone else who worked on theory at some point in their academic career, I have some thoughts on the recent developments in AI and how it may impact the way that we perform and evaluate theoretical research.
We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability.
benchmarks.bespokelabs.ai/autoresearchex…
In each task, we give agents 24 hours
1/6 🧵 Calibration is hard. Multicalibration—fixing errors across every possible subgroup—is usually impossible at scale. Until now. Introducing MCGrad: A production-ready multicalibration library from Meta, accepted at KDD 2026. 🚀 github.com/facebookincuba…
Cursor made me a chrome extension which redirects any html arxiv links that you stumble across on the internet to the pdf version of the arxiv paper instead. Could be useful for some others, but use at your own risk!