My first blog post ever! Be harsh, but, you know, constructive.
Too much efficiency makes everything worse: overfitting and the strong version of Goodhart's law
sohl-dickstein.github.io/2022/11/06/str…
🧵
This is the kind of day that makes me feel good about my life choices. An unusual thing about Anthropic is that we are trying hard on the inside to do what we say we're trying to do on the outside. The conflict with the DoW isn't great, but it is nice when our commitment to our
When AI fails, will it do so by coherently pursuing the wrong goals? Or will it fail the way humans often fail, and take incoherent actions that don't pursue any consistent goal. In other words, like a “hot mess?”
How will this change when AI performing limited tasks transitions
Title: Advice for a young investigator in the first and last days of the Anthropocene
Abstract: Within just a few years, it is likely that we will create AI systems that outperform the best humans on all intellectual tasks. This will have implications for your research and
I will be attending ICML next week. Reach out (by email) if you'd like to chat! About Anthropic / research / life. I'm especially interested in meeting grad students who can teach me new research ideas.