Log inSign up
Arena.ai
3,860 posts
Arena.ai profile banner
@arena

Arena.ai

@arena
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs
US
arena.ai
Joined March 2023
220
Following
223.4K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @arena
    Arena.ai
    @arena
    Jun 4
    Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
    Image
    00:00
    90
  • @arena
    Arena.ai
    @arena
    Sep 6
    Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1! At a $4.14 median cost per task and a +15.8% net improvement, it has reshaped the Pareto frontier! Claude Fable 5.1 is both the most performant and costly among all Agent Arena models today. Based on
    Image
    @arena
    Arena.ai
    @arena
    Sep 6
    Image
    Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. By signal, Claude Fable 5.1 sees
    32
  • @arena
    Arena.ai
    @arena
    Sep 6
    Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. By signal, Claude Fable 5.1 sees
    Image
    @claudeai
    Claude
    Anthropic
    @claudeai
    Sep 1
    Image
    00:27
    We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
    50
  • @arena
    Arena.ai
    @arena
    Sep 5
    Grok Imagine Video 1.5 Agent from @SpaceXAI has landed in the Text-to-Video Arena at #5 with 1491 pts! This release performs on par with Wan-3.0 and FLUX 3 Video (each with 1494 pts), just 3 pts below. Grok Imagine Video 1.5 Agent out performs both Dreamina Seedance-2.5 and 2.0,
    Image
    @grok
    Grok
    SpaceXAI
    @grok
    Sep 5
    Image
    00:11
    Grok Imagine Video 1.5 agent is now available. Powered by our newest Image 2.0 model, it delivers higher quality, better storytelling from a smarter agent and excels at connecting multiple shots together with greater continuity.
    20
  • @arena
    Arena.ai
    @arena
    Sep 5
    GPT-6 Astra (Max) by @OpenAI is #1 on Code Arena: WebDev reshaping the Pareto frontier! It is SOTA performance at $40/Mtoken, matching the latest Claude model pricing. In Code Arena, AI models are ranked on web development tasks, including agentic coding workflows that require
    Image
    @arena
    Arena.ai
    @arena
    Sep 5
    Image
    Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing. GPT-6 Astra by @OpenAI takes the top spot in Code Arena: WebDev with a
    40
Advertisement
Advertisement