4 instrumentors shipped today thanks to our cracked OSS community.
🤖 AG2: multi-agent conversations, group chats, and tool calls
⚡ Together AI: fast inference across OSS models
🧠 Cohere: enterprise chat, search, and RAG
🦙 Ollama: local models galore
Open-Source AI Observability and Evaluation
- AI Queries and Browser AI Phoenix filters now understand English. Type "responses with apologies" and get a real filter expression back: ▎ 'sorry' in output.value or 'apolog' in output.value
- The experiment said the fix works, but experiments aren't production. Trace → dataset → evaluator → experiment → verify in production. PXI helps you every step of the way!
- Balancing cost, speed, and correctness can be a tricky balance. That's why we need good visuals to figure out the right sweet spot!
- Customizable visualizations of your agent traces are here. Track cache hits, online eval degradations, tool call errors, all in real-time so you can quickly identify problems in production. What does your agent operations center look like?

