Trace OpenAI Codex sessions in Braintrust with the `trace-codex` plugin.
Add the plugin to your Codex setup to capture sessions as hierarchical traces. Works in interactive sessions and CI runs, with config options for projects, metadata, and flush behavior.
Read more →
Kimi K3 and DeepSeek V4 Flash are now available as built-in models on Braintrust, joining GLM-5.2. Open models keep improving, and it should be easy to test candidate models against your own prompts, datasets, and production traces.
We ran all three on 327 of the hardest
Eve's legal agents analyze thousands of documents, build medical chronologies, summarize eight-hour depositions, run legal research across case law and statutes, and draft work product for cases.
To keep that quality consistent, they built Plaintiff Bench, a benchmark of 400+
New in the experiments page:
- The analysis chart now includes options for All scores (avg) and Scale by axis, so you can plot duration, cost, or any metric against your scores
- A new summary table that compares scores across all selected experiments and highlights the best and