Log inSign up
Guanghan Ning
fleet
75 posts
Guanghan Ning profile banner
@quietnning

Guanghan Ning

fleet
@quietnning
Research @fleet_ai, formerly @ ByteDance-Seed (LLM) blog.guanghan.ai · views my own
San Francisco, CA
guanghan.ai
Joined July 2015
272
Following
640
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Apr 28
    Personal news: After ByteDance Seed and a stint as an independent researcher, I'm joining @fleet_ai as a Member of Technical Staff on the research team 🚀 Building Witness-inspired puzzle environments for ARC-AGI-3 convinced me that RL environments are one of the most
    Image
    10
  • @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Sep 3
    GPT-6 Astra came out, achieving 62.7% score on the held-out eval set of ARC-AGI-3, a (long-due) huge jump! Wondering how much of this progress came from recurrent/looped transformer architecture (that offers deeper implicit reasoning), how much from the symbolic world modeling
    @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Jun 9
    Excited to see how Fable 5 performs on ARC-AGI-3, and here is my take: 1) The capability jumps on agentic benchmarks often come from the external shell, e.g., a self-referential, code-as-harness scaffold (Darwin Gödel Machine: freeze the model, evolve the scaffolding). However,
    6
  • @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Aug 13
    interesting benchmark! Been building something similar. Would be a great benchmark companion to cross-check models’ capabilities.
    @jcrwhittington
    James Whittington
    @jcrwhittington
    Aug 12
    We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural
    Image
  • @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Jul 24
    Opus 5 reports 30% on ARC-AGI-3, ~4× the previous best model, ~20× its predecessor Opus 4.8. We tested it on Witness, our held-out suite of ARC-AGI-3-style interactive puzzle games. The leap doesn't transfer. On Witness composites (same harness, same budget for every model),
    Image
    @claudeai
    Claude
    Anthropic
    @claudeai
    Jul 24
    Replying to @claudeai
    Image
    On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.
    46
  • @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Jul 19
    Also looking forward to evaluating & taking a glimpse at Qwen3.8-Max-Preview's performance, as soon as it is available on openrouter!
    @quietnning
    Guanghan Ning
    fleet
    @quietnning
    Jul 19
    Kimi-K3 is the real deal! It is at fable-5 level, clearly surpassing Opus4.8. (kimi-k3 demonstrating better stability across seeds; fable-5 has higher ceiling, better Bo5 performance. fable-5 results may include fallback behavior.) Congrats to the @Kimi_Moonshot team on this
    Image
Advertisement
Advertisement