Log inSign up
Braintrust
891 posts
Braintrust profile banner
@braintrust

Braintrust

@braintrust
Active observability for agents in production.
braintrust.dev
Joined August 2023
60
Following
7,555
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @braintrust
    Braintrust
    @braintrust
    Sep 3
    Your team deserves more from your agent observability platform - a single, connected place for instrumentation, investigation, and measurement, enhanced with intelligence. In Braintrust, you can start with an open-ended prompt about agent behavior and carry the investigation
    Image
    00:00
    12
  • @braintrust
    Braintrust
    @braintrust
    Sep 2
    Use the Braintrust JavaScript/TypeScript SDK for tracing and evaling agents in any JS or TS project. Includes integrations for OpenAI Agents, OpenTelemetry, and Temporal. Run evals with a single CLI command or add automatic instrumentation with no code changes. Read more →
    Image
  • @braintrust
    Braintrust
    @braintrust
    Sep 1
    Cloudflare Agents now emits native OpenTelemetry traces, and you can route them straight to Braintrust. Export spans via OTLP, instrument your Workers in JavaScript using the Cloudflare Agents SDK, @cloudflare/ai-chat, @cloudflare/think, or Flue. Turn production traces into
    Image
  • @braintrust
    Braintrust
    @braintrust
    Aug 31
    Imagine doing your job without ever looking anything up on the internet. That's an agent without web search. Giving agents the web changed what they can do, especially on anything recent that isn't reflected in training data. But agents don't search like people, so optimizing
    Image
    5
  • @braintrust
    Braintrust
    @braintrust
    Aug 28
    The Braintrust eval library has a repo of skills your coding agent can read, so you can easily build and run evals on your data. Here's an example of using a skill in Claude Code to compare the Codex CLI and Pi, both running GPT-5.6 Sol, on a 30-task stratified SWE-bench
    Image
    00:00
    2
Advertisement
Advertisement