This is an interesting benchmark, if for no other reason than how “low” the scores seem to be at this point.
We’ve got great scores on ARC-AGI but only the beefiest models can solve feature tickets 1/3rd of the time.
Agents on Rails: Stage 2 is live. We wanted to find out: can you hand a model a real feature ticket and trust what comes back?
The jump from atomic tasks to feature requests has interesting results…
@OpenAI GPT-6 Astra is new to the leaderboard, and it came out on top: 35% of
codex pro tip: you don’t have to stare at the screen while your sessions are running
you can cook while listening to jazz… jump in a cold lake… take yourself to a lovely restaurant… sit on a porch and listen to the leaves rustle in the wind…
Cursed prompt:
I want you to spawn 20 subagents and find the most valuable line in the code base to change. Create exactly one 1-line PR. Take as much time as you need.
@ranjanxroy@Kantrowitz RE: Anthropic and Salesforce.
Claude needs access to Slack for bringing metered usage products to the masses (non-developers). Claude Tag racks up usage quite fast, and is most useful in Slack.