Log inSign up
Wei-Lin Chiang
Arena.ai
832 posts
@infwinston

Wei-Lin Chiang

Arena.ai
@infwinston
Building @Arena, PhD in AI & systems @UCBerkeley
linkedin.com/in/wei-lin-chi…
Joined February 2012
965
Following
5,627
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @infwinston
    Wei-Lin Chiang
    Arena.ai
    @infwinston
    Aug 5
    DeepSeek-v4-Flash on the cost-performance frontier for agentic tasks.
    @arena
    Arena.ai
    @arena
    Aug 5
    DeepSeek-V4-Flash (High) by @deepseek_ai has reshaped the cost-performance Pareto frontier in Agent Arena, with a $0.024 median cost per task! It lands to the right of GPT-5.6 Luna (xHigh) which has a $0.026 median cost per task, and to the left of DeepSeek-V4-Pro (Thinking) at
    Image
    1
  • @infwinston
    Wei-Lin Chiang
    Arena.ai
    @infwinston
    Aug 1
    The latest DeepSeek release is underhyped. DeepSeek-V4-Flash pushes the intelligence/$ frontier forward by an order of magnitude, and it's open-source. This will bring powerful agentic capabilities to the mass market and unlock 10x more use cases overnight.
    @arena
    Arena.ai
    @arena
    Aug 1
    Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with a score of 1586! Priced at $0.14/$0.28 per MToken, it’s the best performance-per-dollar of any model in its class. Congrats to the @deepseek_ai team!
    Image
    17
  • @infwinston
    Wei-Lin Chiang
    Arena.ai
    @infwinston
    Aug 1
    Incredible breakthrough by @deepseek_ai - DeepSeek-V4-Flash reaches Opus-level performance with 30x cheaper cost.
    @arena
    Arena.ai
    @arena
    Aug 1
    DeepSeek-V4-Flash-High by @deepseek_ai is #7 overall in the Frontend Code Arena with 1,586 pts! It’s #3 among open. In categories it’s #4 in Consumer Product, #6 in Reference-based Design, Data & Analytics, Gaming, and #7 in Brand & Marketing. This is an impressive improvement
    Image
    1
  • @infwinston
    Wei-Lin Chiang
    Arena.ai
    @infwinston
    Jul 30
    Three key dimensions to scale frontier AI evals - quality - speed - cost Autoevals give you best of all worlds. Fast, efficient, and high-quality, powered by millions of real-world tasks in Arena.
    @arena
    Arena.ai
    @arena
    Jul 30
    Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences. Highlights: - High-quality evaluation signals calibrated on real preference data - Strong alignment with live human evaluations -
    Image
    1
  • @infwinston
    Wei-Lin Chiang
    Arena.ai
    @infwinston
    Jul 29
    Test Time Scaling in Agent Arena. - Opus-5 test-time scaling beyond GPT-5.6 sol - Opus-5-Max reaching Fable performance but more expensive
    @arena
    Arena.ai
    @arena
    Jul 29
    Claude Opus 5 vs GPT 5.6 in Agent Arena Highlights: - Opus (High) and Opus (Max) outperform GPT Sol (xHigh), but at higher cost. - Opus (Medium) matches GPT Sol (xHigh) in performance at approximately the same cost. - Fable 5 (High) achieves the most optimal price vs.
    Image
    1
Advertisement
Advertisement