Where AI meets the real world. Formerly LMArena. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring → arena.ai/jobs
Introducing Agent Mode: Agentic AI is now measured in the Arena.
Agent Mode can do deep research, create reports, generate images, build websites, debug code, and more.
It completes more complex tasks by using tools like web search, bash in a sandbox environment, image
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
Chinese AI lab Moonshot just beat Claude Fable 5 in coding.
“The Chinese labs might actually just be really good at developing models and not just distilling American intelligence.” — @arena CEO, @ml_angelopoulos
For the first time, China has taken the lead over the US in Frontend Code Arena with the launch of Kimi-K3 by @Kimi_Moonshot.
The last time a Chinese model came close was in early 2025, with DeepSeek-R1.
This may be the single biggest release of the year, and marks the moment that OSS Chinese models have surpassed US models.
On Code Arena, Kimi K3 has BEATEN FABLE.
This is only 6 weeks after the Fable release.
This makes @Kimi_Moonshot the #1 AI lab in the world on frontend
Factuality and human preference are complementary signals.
00:00 Why we're adding factuality as a signal
00:14 Human preference vs. factuality: complementary, not the same
01:06 How this differs from style control
02:32 The math: reviewing the existing Bradley-Terry objective
Introducing factuality in the Arena: a new ranking of models according to a weighted combination of human preference and factuality.
Model rankings are now viewable according to a weighted combination of human preference and factuality. Factuality is live in our Text and Search
Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate.
When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average.
For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50%
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
In the Text Arena, Kimi-K3 by @Kimi_Moonshot landed #9, with 1486 pts.
This is another significant improvement from Kimi-k2.6 (#38 -> #9).
- Top 10 in Creative Writing, Coding and Instruction Following
- #1 in three occupations: Physical & Social Science, Legal & Government,
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,
In the Text Arena, Inkling also landed at #10 for open source (1447 pts), and #58 overall. It is the #2 open source model from the US, second to GoogleDeepMind’s Gemma-4-31B.
Highlights across key categories:
- #10 Hard prompts
- #12 Longer Query
- #14 Coding
- #15 Instruction