Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
We analyzed how @claudeai's writing has changed from Fable 5 to Fable 5.1 across tens of thousands of high-reasoning Text Arena outputs.
Overall, Fable 5.1 by @AnthropicAI uses fewer agreement openers, fewer em dashes and less wording like “honestly” and “frankly”, while its
We analyzed both @claudeai Fable 5 and Fable 5.1 on writing clarity ranked by humans, and are sharing results later today.
Which model do you think writes more clearly?
Game development remains one of the most-requested, and most-challenging, categories on Arena.
@iamwaynechi, PhD candidate at Carnegie Mellon University and research intern at Arena, just walked us through GameDevBench: a benchmark built from real tutorials that turns game
We analyzed both @claudeai Fable 5 and Fable 5.1 on writing clarity ranked by humans, and are sharing results later today.
Which model do you think writes more clearly?
For a limited time, you can test GPT-6 Astra by @OpenAI directly in Arena! Head to Direct Mode, and select it from the dropdown menu. Find the link below.
GPT-6 Astra (Medium) is available in Direct Mode for the next 24 hours, through Sept. 10 at 8am PT. It will remain available