Primarily researched RL for agents, dedicating the next decade
to program synthesis and inference.
Studied at Stanford under Stephen Boyd (convex optimization),
alongside HPC and a handful of other subjects.
Off-screen: soccer, singing, jiu-jitsu, poker.
Selected Work
02
An RL training run on agents playing a Minecraft simulation, done on
PufferLib.
CUDA kernel optimization for an AlphaFold-3 style TriMul kernel, reaching 1185µs latency
live dashboard