1. X
  2. Dan Luu
Log inSign up
Dan Luu
5,582 posts
user avatar
Dan Luu
@danluu
Active on mastodon.social/@danluu; also trying out bsky.app/profile/da. No longer read replies or notifications here now that tweetdeck is gated.
danluu.com
Joined December 2008
43
Following
46.2K
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Dan Luu
    @danluu
    Dec 24, 2017
    Computer latency: 1977-2017 danluu.com/input-lag/
    Image
  • user avatar
    Dan Luu
    @danluu
    Jul 24
    Exercises in evals and benchmarking: DeepSWE / Senior SWE-Bench, performance math, and cold weather tires danluu.com/exercise-7/
    Image
    Image
    Image
    Image
  • user avatar
    Dan Luu
    @danluu
    Jul 3
    Agentic coding notes from Galapogos Island: danluu.com/ai-coding/
    I've been using AI fairly heavily since last November and the whole thing is a funny experience. An agent will do something that, if a human did it, you'd immediately fire them. My reaction, of course, is to act as if this is great and spin up a thousand agents so they can do even more of that.

Mid-last year, I had GPT (maybe 5.0 or 5.1) try to find the source of a bug. Naturally, this code didn't have tests and git bisect wouldn't work, and it was a UI interaction bug for which I'm not even really qualified to write a test for, so I asked codex to bisect between dates X and Y to find the commit that introduced this bug. Codex immediately told me the offending commit was after this date range (which couldn't possibly be correct). On telling codex this was wrong, it then told me some commit that was obviously also not the offending commit once or twice. On telling it those were wrong, it then told me the offending commit was some plausible looking commit. When I asked it to prove or...
  • user avatar
    Dan Luu
    @danluu
    Jun 8
    Exercises in benchmarking, evals, and experimental design, part 6: patreon.com/posts/160473058
    26. There's a quote being passed around that reads

Steve Yegge says people using AI coding agents are "10x to 100x as productive as engineers using Cursor and chat today, and roughly 1000x as productive as Googlers were back in 2005."
That's a real number. I've seen it. I've lived it.
What's wrong with this benchmark?

27. At work, people keep recommending Caveman Mode for reducing token usage and speeding up tasks. The README in the repo claims a 3x speedup on tasks with a 75% reduction in token usage. On HN, people say this approach can't work and will make things worse for theoretical reasons that are guaranteed to hold, such as https://news.ycombinator.com/item?id=47647907. Someone tried benchmarking caveman mode and found https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words. The world's biggest programming influencer said, "it actually works; it actually works quite well ... no, I'm not exaggerating".

What's wrong with these benchmarks?

28. In a viral tw...
  • user avatar
    Dan Luu
    @danluu
    May 29
    In 10 years, will at least one of {Anthropic, OpenAI} be a public company valued at >= $1T in May 2026 dollars after adjusting for S&P 500 valuations (analogous to inflation adjustment, but using S&P 500 instead of CPI)?
    Yes77.2%
    No22.8%
    92 votesFinal results

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement