I build sane open-source RL tools. MIT PhD, creator of Neural MMO and founder of PufferAI. DM for business: non-LLM sim engineering, RL R&D, infra & support.
The world runs on TypeScript & JavaScript.
Our bet is that AI engineering will follow suit. The growth in @aisdk downloads and adoption has been astonishing.
When we wrote the Ship AI keynote it was at 3.4M weekly downloads. A couple weeks later, it’s now at 4.1M 😳
You have to have been in ML for over a decade to really understand how bad this was. From 2015-2018ish, any time anything went wrong in an experiment, at least someone would say "local minimum"
I don't care if Karpathy is down on RL. He, Carmack, Ilya, and Alec Radford could all show up in person to tell me I'm wasting my time and I'd still keep doing it. Because damn it this tech is too cool not to exist.
RL really sucks. It takes 10 hours just to learn breakout.
... a few years ago. It's <30 seconds on 1 GPU now in PufferLib and still dropping. Write faster code.
PufferLib 3.0: We trained reinforcement learning agents on 1 Petabyte / 12,000 years of data with 1 server. Now you can, too! Our latest release includes algorithmic breakthroughs, massively faster training, and 10 new environments. Live demos on our site. Volume on for trailer!