1. X
  2. Sentient
Log inSign up
Sentient
2,496 posts
Image
user avatar
Sentient
@SentientAGI
To ensure that Artificial General Intelligence is open-source and not controlled by any single entity. @SentientEco @OpenAGISummit
San Francisco, CA
sentient.xyz
Joined February 2024
69
Following
535.7K
Followers
AffiliatesAffiliatesRepliesRepliesArticlesArticlesMediaMedia
  • Pinned
    user avatar
    Sentient
    @SentientAGI
    Aug 4
    Agents can fetch data and nail the facts, but combining it all into real research is where they start to break. Sentient researchers Darshan Tank (@TankDarshan7) and Sidhant Rahi tackle this in "CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis", accepted into
    Image
  • user avatar
    Sentient
    @SentientAGI
    22m
    This is why checking your agent's math isn't enough ↓
    user avatar
    Sentient Ecosystem
    @SentientEco
    23m
    Most AI agents don't just get the math wrong. They get the number from the wrong place: outdated table, mismatched row, wrong year. That's why @abraxasnz13 built an agent that won't calculate until it knows where the number came from ↓
    Image
    00:00
  • user avatar
    Sentient
    @SentientAGI
    Aug 7
    Proof that a better harness won't crack a harder task ↓
    user avatar
    Sentient Ecosystem
    @SentientEco
    Aug 7
    You can't benchmark your way around a hard problem. Across 13,000+ agent runs in the OfficeQA Public Challenge, Sentient researchers @iamnamanvats and Deep Halder found that different harnesses agreed 88-93% of the time on which tasks succeeded and which failed. TLDR:
    Image
  • user avatar
    Sentient
    @SentientAGI
    Aug 6
    Turns out our runner-up's best piece of prompt engineering was a backspace ↓
    user avatar
    Sentient Ecosystem
    @SentientEco
    Aug 6
    One of the best decisions @Fr0oZi made during the OfficeQA Public Challenge was to delete a single line: "You are a financial analyst." Here's how he placed 2nd in the Arena using @goose_oss and @MiniMax_AI M2.7 ↓
    Image
    00:00
  • user avatar
    Sentient
    @SentientAGI
    Aug 5
    More skills should mean smarter agents right? Sentient researchers Darshan Tank (@TankDarshan7) and Baran Nama found the opposite happens: adding new skills can break tasks the agent has already solved. They call this the "Regression Tax", a hidden cost that erased 59% of the
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement