Log inSign up
Ofir Press
3,443 posts
Ofir Press profile banner
@OfirPress

Ofir Press

@OfirPress
I push the AI frontier by building tough benchmarks with amazing people. SWE-bench, SWE-agent, SciCode, AlgoTune. Postdoc @Princeton. PhD @nlpnoah @UW.
NYC
ofir.io/about
Joined June 2016
9,343
Following
19.5K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @OfirPress
    Ofir Press
    @OfirPress
    May 5
    1) Our team at Meta has a tough new coding benchmark challenging models to code entire programs including ffmpeg and the PHP compiler from scratch. 2) Top accuracy is 0% 3) We will be making the benchmark harder.
    @jyangballin
    John Yang
    @jyangballin
    May 5
    How much of SQLite, FFmpeg, PHP compiler can LMs code from scratch? Given just an executable and no starter code or internet access. Introducing ProgramBench: 200 rigorous, whole-repo generation tasks where models design, build, and ship a working program end to end. 🧵
    Image
    37
  • @OfirPress
    Ofir Press
    @OfirPress
    14h
    As of a few weeks ago I can officially say that I have 10 years of language modeling experience. My first paper showed that you could use the same matrix as both the embedding and [what's now known as the] unembedding layer. I started working on it in Jan 2016, submitted it to
    Image
    4
  • @OfirPress
    Ofir Press
    @OfirPress
    Sep 5
    i had astra recreate the arch at washington square park in nyc. its not perfect but this is 100x better than even six months ago. super cool
    Image
    2
  • @OfirPress
    Ofir Press
    @OfirPress
    Sep 3
    interesting word choice in the astra report
    Image
    5
  • @OfirPress
    Ofir Press
    @OfirPress
    Sep 1
    SWE-bench Multimodal was used in today's Anthropic launch. No model passes the 60% mark. Lots of room for growth
    Image
    @jyangballin
    John Yang
    @jyangballin
    Sep 1
    Image
    Releasing SWE-bench Multimodal v2.0 today 480 tasks where a coding agent must interpret visual assets like screenshots, diagrams, recordings to diagnose and fix a bug in a repository.
    7
Advertisement
Advertisement