1. X
  2. Kilian Lieret
Log inSign up
Kilian Lieret
592 posts
user avatar

Kilian Lieret

@KLieret
Meta Superintelligence, prev. Princeton. SWE-bench multilingual/multimodal, SWE-agent, mini-swe-agent, SWE-smith, CodeClash, ProgramBench
New York
github.com/klieret
Joined May 2021
253
Following
2,569
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Kilian Lieret
    @KLieret
    May 12
    The first ProgramBench task was just solved by GPT 5.5 high/xhigh. Interestingly, high/xhigh picked two different languages for the task (C vs Python). GPT 5.5 xhigh was significantly better than Opus 4.7 xhigh in all metrics. 🧵
    Image
  • user avatar
    Kilian Lieret
    @KLieret
    Aug 17
    mini-swe-agent with Opus 5 is significantly ahead on TerminalBench 3/FrontierBench. Interestingly it even beats Fable 5 + Claude Code.
    Image
  • user avatar
    Kilian Lieret
    @KLieret
    Aug 14
    Had a great time talking at @ScaleAILabs. New levels of model capabilities require new benchmarking paradigms. Also really enjoyed talks by @sirrice @mohit_r9a and @MiguelR33478246! Thanks to all the organizers and to @liuying04
    Image
    Image
  • user avatar
    Kilian Lieret
    @KLieret
    Aug 13
    Extremely cool project: What's the best SWE-bench score you can achieve by training a model from scratch with limited $ budget. SWE-bench might be near saturated for frontier models, but it still has a lot of life in it for hillclimbing with smaller or from scratch models
    user avatar
    Ricardo Olmedo
    @rdolmedo_
    Aug 13
    We turned @karpathy's nanochat into a SWE-bench speedrun! ⚔ $60 of train compute → 5.0% pass@1 ā‰ˆ Claude 2 (2023 SOTA) ⚔ $1000 → 11.0% pass@1 > Claude 3 Haiku From randomly initialized model weights! How? šŸ§µšŸ‘‡
    Image
  • user avatar
    Kilian Lieret
    @KLieret
    Aug 13
    When SWE-bench was launched initially, people questioned whether it was feasible at all. We got very similar feedback for ProgramBench. Up until now, only 2/200 tasks were solved. But Opus 5 just made an enormous jump with 9 fully solved tasks instances!
    user avatar
    John Yang
    @jyangballin
    Aug 13
    Takeoff fully in motion: Claude Opus 5 (xhigh) is the new #1 on ProgramBench, and it's not close. ProgramBench asks a coding agent to rebuild a whole program (sqlite, ffmpeg, php) from scratch. Previous high: GPT 5.6 Sol w/ 2 Opus fully resolves *9* (= 4.5%)
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsĀ·PrivacyĀ·CookiesĀ·AccessibilityĀ·Ads InfoĀ·Ā© 2026 X Corp.
Advertisement
Advertisement