Accounting AI doing real work. Join us getbasis.ai
- To trust any long-horizon agent, you have to supervise how it works over a full trajectory, rather than just looking at its final answer. Behavior specs are now an open standard, built with @braintrust and based on how we actually evaluate our production agents:Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days. Even if you




