Where AI meets the real world. Formerly LMArena. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs
Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
MiniMax H3 by @MiniMax_AI is in the Video Arena!
Available in Text-To-Video and Image-To-Video arenas. Open weights coming soon.
Bring your most creative prompts and start voting. Scores on the way.
MiniMax H3: Omni-Reference, Commercial-Grade Generation, Unbeatable Cost Efficiency, Open Weights
Your creative destiny, on your terms.
Now Live at HailuoAI.video & MiniMax API.
Big update: @OpenAI has decreased the price of GPT-5.6 Luna by 80% and Terra by 20%.
Luna is now priced at $0.20/$1.20, and has strong gains from additional reasoning while remaining at an efficient cost. Terra is priced at $2/$12.
This is frontier-level intelligence at a
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are
Inkling-Small (@thinkymachines) debuts at rank ~#88 (1431 pts, AutoEval) in Text Arena and ~#21 among open.
This makes it one of only five US models in the top 25 open models overall.
Note: this is an early AutoEval score, in which a Reward Model trained on Arena's human
Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
thinkingmachines.ai/news/inkling-s…
Fine-tune it on Tinker today, or chat with
Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences.
Highlights:
- High-quality evaluation signals calibrated on real preference data
- Strong alignment with live human evaluations
-