1. X
  2. Arena.ai
Log inSign up
Arena.ai
3,523 posts
Image
user avatar
Arena.ai
@arena
Where AI meets the real world. Formerly LMArena. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs
US
arena.ai
Joined March 2023
217
Following
199.9K
Followers
AffiliatesAffiliatesRepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • Pinned
    user avatar
    Arena.ai
    @arena
    Jun 4
    Introducing Agent Mode: Agentic AI is now measured in the Arena. Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more. It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
    Image
    00:00
    353K0353K
  • user avatar
    Arena.ai
    @arena
    Jul 17
    Big news from @Kimi_Moonshot this week. Check out Kimi K3 head-to-head with Fable 5 on identical prompts in Frontend Code Arena.
    Image
    00:00
    Image
    user avatar
    Arena.ai
    @arena
    Jul 16
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
    115K0115K
    user avatar
    Arena.ai
    @arena
    Jul 17
    Let us know what you think, and visit arena.ai/code to generate and share your own.
    8.7K08.7K
    user avatar
    Arena.ai
    @arena
    Jul 17
    0:00 Ramen Service Night Website - Kimi K3: …-b853-81a7741caf77.staging.arena.site - Fable 5: …-a869-a94b040a5112.staging.arena.site 0:06 Deep-Sea Research Station Dashboard - Kimi K3: …-bde7-55e3188be393.staging.arena.site - Fable 5: …-a25b-7c20e4310ca8.staging.arena.site 0:12 Theater Stage Rehearsal Site - Kimi K3:
    Image
    Maru Menya — Neighborhood Ramen, Live Friday Night
    From 019f6c6f-96be-76d9-b853-81a7741caf77.staging.arena.site
    8.5K08.5K
  • Arena.ai reposted
    user avatar
    The Information
    @theinformation
    Jul 17
    Chinese AI lab Moonshot just beat Claude Fable 5 in coding. “The Chinese labs might actually just be really good at developing models and not just distilling American intelligence.” — @arena CEO, @ml_angelopoulos
    Image
    00:00
    39K039K
  • user avatar
    Arena.ai
    @arena
    Jul 17
    Replying to @arena
    See the full Frontend Code Arena leaderboards at arena.ai/leaderboard/co…
    13K013K
  • user avatar
    Arena.ai
    @arena
    Jul 17
    For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot. The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
    Image
    Image
    user avatar
    Anastasios Nikolas Angelopoulos
    Arena.ai
    @ml_angelopoulos
    Jul 16
    This may be the single biggest release of the year, and marks the moment that OSS Chinese models have surpassed US models. On Code Arena, Kimi K3 has BEATEN FABLE. This is only 6 weeks after the Fable release. This makes @Kimi_Moonshot the #1 AI lab in the world on frontend
    450K0450K
  • user avatar
    Arena.ai
    @arena
    Jul 17
    Factuality and human preference are complementary signals. 00:00 Why we're adding factuality as a signal 00:14 Human preference vs. factuality: complementary, not the same 01:06 How this differs from style control 02:32 The math: reviewing the existing Bradley-Terry objective
    Image
    00:00
    Image
    user avatar
    Arena.ai
    @arena
    Jul 15
    Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality. Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search
    32K032K
    user avatar
    Arena.ai
    @arena
    Jul 17
    Subscribe to the Arena YouTube channel for more researcher deep dives: youtube.com/@ArenaAIOffici…
    5.5K05.5K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate. When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average. For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50%
    Image
    Image
    user avatar
    Arena.ai
    @arena
    Jul 16
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
    374K0374K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    In the Text Arena, Kimi-K3 by @Kimi_Moonshot landed #9, with 1486 pts. This is another significant improvement from Kimi-k2.6 (#38 -> #9). - Top 10 in Creative Writing, Coding and Instruction Following - #1 in three occupations: Physical & Social Science, Legal & Government,
    Image
    Image
    Image
    user avatar
    Kimi.ai
    @Kimi_Moonshot
    Jul 16
    Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
    87K087K
    user avatar
    Arena.ai
    @arena
    Jul 16
    See the full Text Arena leaderboards at arena.ai/leaderboard/te…
    12K012K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Kimi K3 from @Kimi_Moonshot has moved the Pareto Frontier for Code Arena: Frontend.
    Image
    Image
    user avatar
    Arena.ai
    @arena
    Jul 16
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
    89K089K
    user avatar
    Arena.ai
    @arena
    Jul 16
    Check out the Pareto Frontier for Code Arena: Frontend
    Image
    WebDev AI Leaderboard - Best AI Models for Web Development
    From arena.ai
    11K011K
  • Arena.ai reposted
    user avatar
    Kimi.ai
    @Kimi_Moonshot
    Jul 16
    🤯
    user avatar
    Arena.ai
    @arena
    Jul 16
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
    Image
    792K0792K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Replying to @arena
    Check out @Kimi_Moonshot's Kimi-K3 official release:
    user avatar
    Kimi.ai
    @Kimi_Moonshot
    Jul 16
    Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
    Image
    Image
    316K0316K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
    Image
    Image
    00:56
    user avatar
    Kimi.ai
    @Kimi_Moonshot
    Jul 16
    Meet Kimi K3
    19M019M
    user avatar
    Arena.ai
    @arena
    Jul 16
    See the full Frontend Code Arena leaderboards at arena.ai/leaderboard/co…
    243K0243K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Replying to @arena
    Dive into the Inkling scores across both leaderboards at: arena.ai/leaderboard
    5.4K05.4K
  • user avatar
    Arena.ai
    @arena
    Jul 16
    Replying to @arena
    In the Text Arena, Inkling also landed at #10 for open source (1447 pts), and #58 overall. It is the #2 open source model from the US, second to GoogleDeepMind’s Gemma-4-31B. Highlights across key categories: - #10 Hard prompts - #12 Longer Query - #14 Coding - #15 Instruction
    Image
    6.6K06.6K
Advertisement
Advertisement