Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵
TPU Inference Externalization
Full Steam Ahead - InferenceX,
Up to 50% Better Performance per Dollar,
Rapid Externalization of TPU stack,
Growing Customer Base,
Ironwood, TPUv8i, Reducing CUDA Moat
ALERT🚨: AMD INCREASED @vllm_project PERFORMANCE BY 11x IN LESS THAN 19 DAYS ON MI355X AGENTIC WORKLOADS on the modern MiniMax M3 model!
This was done entirely through software optimizations, mainly by optimizing the long-context attention op, along with other optimizations!