Pinned
vllm.cpp runs @MiniMax_AI 's MiniMax-H3 now. 33.1B, video and audio out of a single model, and I'm still a bit stunned that it works
we reimplemented both VAEs from the checkpoint's remote python, so there's no torch anywhere in the process. And you drive the whole thing over




