Introducing Factory Benchmarks: The first model bench generated from your own coding tasks.
Measure, test and improve coding agents by replaying past agent runs, and cut cost-per-PR by 63%+
Here’s how it works 🧵
I’m excited to finally announce the newest edition my Stanford course 𝗧𝗵𝗲 𝗠𝗼𝗱𝗲𝗿𝗻 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗗𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿. It has been 9 months in the making.
Last November, with the release of Claude Opus 4.5, coding agents experienced a step function improvement in
Factory benchmarks let you build your own model bench using your past coding agent runs.
- Mirrors your environment, secrets, and MCPs for each run
- Scores output on judging criteria you define
- Generates a report with model recs using cost vs. quality
Here's how it works: