Pinned
Better empirical methods for deep learning. PhD at @nyuniversity (@CILVRatNYU). Advised by @kchonyc and @hhexiy. Prev: @allen_ai.
I build things. 🤖
- A great idea by @sea_snell: Use finetuning to predict where zero-shot capabilities emerge. This lets you experiment at a smaller scale. The more finetuning data you have, the smaller of a model you can use. Here's how I think about it: a one-time cost collecting data saves youCan we predict emergent capabilities in GPT-N+1🌌 using only GPT-N model checkpoints, which have random performance on the task? We propose a method for doing exactly this in our paper “Predicting Emergent Capabilities by Finetuning”🧵
- Anthropic put out a great primer on statistical methods for LLM evals by @EvMill. Check out his blog too! He's written gems on A/B testing and other topics---just make sure you don't mind losing an afternoon like I did when I first came across it! 😆 evanmiller.orgNew Anthropic research: Adding Error Bars to Evals. AI model evaluations don’t usually include statistics or uncertainty. We think they should. Read the blog post here: anthropic.com/research/stati…
- If scaling no longer makes economic sense, what does that mean for research?? Will we see more work on architecture and fundamentals again? Or, will the current spread of topics remain unchanged? 🤔OpenAI, Google and Anthropic Are Struggling to Build More Advanced AI “Three of the leading artificial intelligence companies are seeing diminishing returns from their costly efforts to develop newer models.” Scaling laws are breaking down economically. bloomberg.com/news/articles/…
- ✨Proud to announce opda v0.7.0 just released! 🥳🎉 The big drop is a new method to fit the noisy quadratic distribution---the probability distribution that determines what you get from random search! 🧵 1/3




