1. X
  2. Labelbox
Log inSign up
Labelbox
284 posts
Image
user avatar
Labelbox
@labelbox
Frontier RL data for the world’s leading AI teams.
San Francisco, CA
labelbox.com
Joined January 2018
147
Following
3,539
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Labelbox
    @labelbox
    Jul 16
    Manu (@manuaero) joined Jason Calacanis on This Week in AI alongside @sarahookr (Adaption Labs) and @spirosx (Resolve AI) to break down the rapidly shifting AI landscape. From the Token Price Wars and Demis Hassabis’ recent call for an AI oversight board, to the Uber vs. Waymo
    Image
    78K
  • user avatar
    Labelbox
    @labelbox
    Jul 7
    Do AI models become less honest when they self-report their own misbehavior? Our Applied Research team studied a key question in chain-of-thought (CoT) monitorability: do models faithfully report their own misbehavior when they are asked to? We introduce Monitorability
    396K
  • user avatar
    Labelbox
    @labelbox
    Jun 25
    1/ Today, we’re introducing Recursion: the RL platform for building, evaluating, and deploying specialist agents. The next phase of AI won’t just be about smarter models. It will be about systems that learn from the unique expertise of every organization. 🧵
    8.7M
  • user avatar
    Labelbox
    @labelbox
    Jun 16
    Where do models change their minds? Natural Language Autoencoders (NLAs) offer a promising way to translate a model’s internal representations into natural language. But the harder question is: where do meaningful decisions actually happen? We tested a workflow for finding
    Image
    3.9M
  • user avatar
    Labelbox
    @labelbox
    May 20
    When AI benchmarks saturate, what comes next? Historically, leaderboard saturation leads to two paths: hyper-specialized questions or increasingly abstract puzzles. A new paper from @Meta Superintelligence Labs introduces a third path: GIM (Grounded Integration Measure).
    When benchmarks saturate, what comes next? Meta’s GIM pushes AI evaluation toward integrated reasoning
    When benchmarks saturate, what comes next? Meta’s GIM pushes AI evaluation toward integrated...
    From labelbox.com
    11M
  • See @labelbox's full profile

    Sign up
    Log in
Advertisement
Advertisement