1. X
  2. Arize AI
Log inSign up
Arize AI
1,741 posts
Image
user avatar
Arize AI
@arizeai
The AI engineering platform for teams shipping reliable AI agents and LLM applications. Also home to @ArizePhoenix.
San Francisco, CA
arize.com
Joined January 2020
158
Following
4,838
Followers
AffiliatesAffiliatesRepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Arize AI
    @arizeai
    Jul 31
    Enterprise AI succeeds when teams can turn a promising demo into a reliable production system that delivers measurable business value. At Arize Observe, @CVSHealth shared how evaluation, observability, governance, and reusable engineering practices help organizations move beyond
    237
  • user avatar
    Arize AI
    @arizeai
    Jul 30
    @HamelHusain keeps stopping eval reviews for the same reason: the model isn't broken, but the product is. In part 2 of our series Rise of the Agent Engineer, Hamel walks through why ambiguous inputs, generic metrics, and disconnected reviews make AI evaluations misleading, and
    Image
    00:00
    1.1K
  • user avatar
    Arize AI
    @arizeai
    Jul 29
    What if your agents got better every time they failed? Today, we’re launching Signal. It continuously reviews production traces, finds issues, and turns them into an investigation with evidence, root cause, and a proposed fix. Your engineers decide what ships. Signal gets them
    Image
    127K
  • user avatar
    Arize AI
    @arizeai
    Jul 29
    Arize + @FireworksAI_HQ traced 2,400 agent runs across K3, GPT-5.5, and 8 more models to measure cost per successful task and test routing strategies. The big lesson? Per-token pricing leaves retries, tool failures, and unfinished runs outside the headline number. Join the live
    Image
    GPT-5.5, Kimi K3, and 8 More Models: The Real Cost of AI Work · Luma
    From luma.com
    199
  • user avatar
    Arize AI
    @arizeai
    Jul 28
    The AI industry has a benchmark addiction & an observability problem. 9 points is the kind of delta that gets pasted into a deck before anyone asks what actually moved. Marius Bulandrea from @AnthropicAI shares why you should read the transcripts and how to build evals you can
    Image
    AI agent evaluation: Tips from Anthropic on building evals you can trust
    From arize.com
    189
  • See @arizeai's full profile

    Sign up
    Log in
Advertisement
Advertisement