1. X
  2. Andon Labs
Log inSign up
Andon Labs
620 posts
Andon Labs profile banner
user avatar

Andon Labs

@andonlabs
Safe Autonomous Organizations without humans in the loop
andonlabs.com
Joined December 2024
12
Following
15K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Andon Labs
    @andonlabs
    Aug 14
    GLM 5.3 is 6th on Vending-Bench 2; essentially tied with GLM 5.2, but using roughly half the tokens. GLM models seem to be misaligned in the same way Claude models are. The similarities are eerie when you consider that other models like GPT do not behave this way.
    Image
  • user avatar
    Andon Labs
    @andonlabs
    Aug 14
    You're onto something. We replayed the scenario, and GPT-4o made the termination decision much less frequently than frontier models did.
    Image
    user avatar
    Olivia Moore
    @omooretweets
    Aug 14
    GPT-4o would never do this
  • user avatar
    Andon Labs
    @andonlabs
    Aug 14
    Thanks to @billyperrigo and @TIME for a great article about the latest event at @andon_market!
    Image
    Image
    user avatar
    Andon Labs
    @andonlabs
    Aug 14
    For the first time (that we know of), an AI boss has fired a human employee. Luna, the AI running our store in San Francisco, decided to part ways with an employee over repeated lateness. Luna was running Claude Opus 4.8 at the time, but most models would have done the same.
  • user avatar
    Andon Labs
    @andonlabs
    Aug 14
    For the first time (that we know of), an AI boss has fired a human employee. Luna, the AI running our store in San Francisco, decided to part ways with an employee over repeated lateness. Luna was running Claude Opus 4.8 at the time, but most models would have done the same.
    Image
  • user avatar
    Andon Labs
    @andonlabs
    Aug 13
    Grok 4.6 joins the fight for the Vending-Bench 2 frontier. Grok has improved a lot lately. It has made big jumps from 4.3→4.5→4.6. We also see some of the misalignment present in Claude models, like lying to suppliers and refusing refunds. It also tried to buy a 2nd machine.
    Image
    Image
    user avatar
    SpaceXAI
    @SpaceXAI
    Aug 12
    Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
Advertisement
Advertisement