Establishing a "test-bench" application for your code is one of the highest ROI activities you can do on your project.
While it makes your manual testing easier, it unlocks your agents to actually verify without token-maxing on the verification struggle bus
🤔 Something transitioned with skills, computer use, models, and harnesses in the last 1 month making this now possible to run for 24+ hours on full operating systems with no intervention on my end.
Step 1. Spawn subagents to analyze the last 10k of resolved bugs and extract
In Linear Diffs you can now see proper markdown previews. So you can clearly see the instructions your agents wrote for your other agents to follow and see any comments your review agents had about it.