Customizable visualizations of your agent traces are here. Track cache hits, online eval degradations, tool call errors, all in real-time so you can quickly identify problems in production.
What does your agent operations center look like?
In an agent session, a cache miss re-bills your entire history at full input price. That's why a "continue" after a coffee break can cost more than the model's actual answer.
Cache read vs. write isn't a footnote in your bill. Earendril's post is a must read.
You can use PXI to run an experiment directly from Phoenix! Here's one that tests the system prompt vs. schema-aware prompt, same model, graded by a code evaluator — no LLM judge needed when the check is programmatic.
TIL: "When an eval fails everything, suspect the eval first."