Log inSign up
penlu
Cognition
708 posts
@penlume

penlu

Cognition
@penlume
human computer interface. naturally occurring feature of your environment
penlu.me
Joined February 2015
1,579
Following
7,821
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @penlume
    penlu
    Cognition
    @penlume
    Sep 3
    4397328654844826923795068102505872571721883526553349659561256924505973939597593482272505698004801207988043088656411102133523080581 divides RSA-260
    614
  • @penlume
    penlu
    Cognition
    @penlume
    Aug 6
    human reward function very sophisticated. like imagine winning the fields medal. this is not the same as taking a hit of opium in many ways
    5
  • @penlume
    penlu
    Cognition
    @penlume
    Jul 27
    even before the computer, lots of people were already engaged in the dissemination and deployment of tech that they did not understand and that was created by others. probably it is even more stark now, and soon
    1
  • @penlume
    penlu
    Cognition
    @penlume
    Jul 24
    among the LLM behaviors, waytoomany LoC output hurts me most I am not aware of SWE benchmarks other than FrontierCode wherein writing more code will eventually reduce your score
    @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Jul 24
    We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria
    Image
    1
  • @penlume
    penlu
    Cognition
    @penlume
    Jul 8
    made with carefully selected ingredients, expertly prepared by our chefs using traditional techniques
    @cognition
    Cognition
    @cognition
    Jul 8
    Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale
    Image
    3
Advertisement
Advertisement