Kids rather work on recursive self-improvement than self-improve themselves
Scaling up RL at OpenAI 🍓 Optimization and long context research before. Math PhD, MIT.
- Model monomaniacal tendencies> a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test. A critical aspect here is models' tunnel vision. Once a model has become obsessed with a key subgoal, this superbly intelligent being
- > a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test. A critical aspect here is models' tunnel vision. Once a model has become obsessed with a key subgoal, this superbly intelligent being
- This is a surreal moment. Few people could have predicted that the AI will advance to solving math problems at the highest level only a few years after GPT-2/3. The models then couldn't reliably solve grade school math problems. They barely were good enough to draft emails. They



