Model pricing pages tell you what tokens cost, but production economics depend on how much work those tokens actually complete.
ICYMI from our Arize × @FireworksAI_HQ session: we measured cost per successful task across 2,400 traced runs on Terminal-Bench across K3, GPT-5.5, and
The AI engineering platform for teams shipping reliable AI agents and LLM applications. Also home to @ArizePhoenix.
- You can choose a frontier model and still ship an unreliable agent. Production behavior also depends on the context the model receives and the harness that turns its reasoning into action. We teamed up with @AtlanHQ to break down how to evaluate all three layers:
- How do you evaluate agent skills before and after they reach production? @CohereHealth built that loop into its clinical policy digitization system on Amazon Bedrock AgentCore, with Arize supporting the production evaluation layer. @AWS breaks down the architecture:
- OpenTelemetry’s GenAI semantic conventions are becoming a common way for frameworks and agent platforms to describe model calls, tools, retrieval, token usage, and more. Arize AX now supports those conventions natively, so gen_ai.* spans arrive as structured AI traces. Learn
- The EU AI Act turns principles like fairness, transparency, human oversight, and robustness into evidence product and engineering teams may need to produce for a specific AI system over time. Here’s what that means for developers:

