Log inSign up
Cheng-Yuan (Sam) Lee
Cognition
231 posts
Cheng-Yuan (Sam) Lee profile banner
@cl571128

Cheng-Yuan (Sam) Lee

Cognition
@cl571128
Research @cognition | prev intern @windsurf | 2x ICPC World Finalist | from Taiwan 馃嚬馃嚰
College Park, MD
Joined July 2023
437
Following
783
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email
Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.
  • @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Sep 3
    In our early testings, we found that GPT-6 Astra sometimes creates extra test files and ignores codebase conventions. But it follows instructions so well that adding a few sentences to the system prompt fixed it.
    Image
    @cognition
    Cognition
    @cognition
    Sep 3
    Image
    GPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
    11
  • @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Sep 2
    We have updated FrontierCode leaderboard to include more models! Check it out! frontiercode.com
    Image
    9
  • @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Aug 19
    Since I started using Devin a lot late last year, I haven鈥檛 been able to go back. The interaction just feels natural, and it makes getting things done so easy. Do it all with Devin!
    @ScottWu46
    Scott Wu
    Cognition
    @ScottWu46
    Aug 19
    At Cognition, one of our values is to go for it all: when faced with a tradeoff, pick the ambition-maximizing direction. When we launched Devin in 2024, we envisioned a future where every team had an infinite army of junior engineers. We were early and Devin wasn't good enough.
    Image
    Image
  • @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Jul 24
    We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria
    Image
    27
  • @cl571128
    Cheng-Yuan (Sam) Lee
    Cognition
    @cl571128
    Jul 8
    One day, all of the research team spent hours in a room together manually solving the RL tasks we used to evaluate our models. I remember solving one of the tasks and realized that the tests are not even testing what the agent was asked to do. Since then, data has become one of
    @cognition
    Cognition
    @cognition
    Jul 8
    Introducing SWE-1.7, the most capable model we鈥檝e trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale
    Image
    16
Advertisement
Advertisement