Palo Alto, CA
Joined March 2026
- I bet they used BF16-throughput as the denominator when training in FP8 or something. By that algebra, I can get you 150% MFU in no time😅. For reference, as far as I know the SOTA Hopper GEMM kernel is ~84% utilization. arxiv.org/abs/2605.05331
- A common misconception is that video models are too slow to run as closed-loop policies. We’ve shown that not only can they be fast, but they’re fast enough to run on a single RTX 5090! It turns out that if you co-design your model architecture and inference optimizations
- In-context learning is such an elegant application of a generative video model as robot policy.




