1. X
  2. arize-phoenix
Log inSign up
arize-phoenix
933 posts
Image
user avatar
arize-phoenix
@ArizePhoenix
Open-Source AI Observability and Evaluation
Notebook or Container
github.com/Arize-ai/phoen…
Joined February 2023
315
Following
1,724
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    arize-phoenix
    @ArizePhoenix
    Aug 1
    Balancing cost, speed, and correctness can be a tricky balance. That's why we need good visuals to figure out the right sweet spot!
    Image
    00:00
    Image
    100
  • user avatar
    arize-phoenix
    @ArizePhoenix
    Jul 30
    Customizable visualizations of your agent traces are here. Track cache hits, online eval degradations, tool call errors, all in real-time so you can quickly identify problems in production. What does your agent operations center look like?
    Image
    00:00
    177
  • user avatar
    arize-phoenix
    @ArizePhoenix
    Jul 29
    New Sandbox in Phoenix - Monty (Python)!
    Image
    Image
    133
  • user avatar
    arize-phoenix
    @ArizePhoenix
    Jul 28
    In an agent session, a cache miss re-bills your entire history at full input price. That's why a "continue" after a coffee break can cost more than the model's actual answer. Cache read vs. write isn't a footnote in your bill. Earendril's post is a must read.
    Image
    138
  • user avatar
    arize-phoenix
    @ArizePhoenix
    Jul 27
    You can use PXI to run an experiment directly from Phoenix! Here's one that tests the system prompt vs. schema-aware prompt, same model, graded by a code evaluator — no LLM judge needed when the check is programmatic. TIL: "When an eval fails everything, suspect the eval first."
    Image
    00:00
    142
  • See @ArizePhoenix's full profile

    Sign up
    Log in
Advertisement
Advertisement