Models in 2026:
All accusations, no fingerprints.
Mark the model, end the argument.
Joined July 2024
- When an agent takes more actions but makes no real progress, this is a useful early-warning sign that it’s heading towards failure. Across 13K+ OfficeQA runs, agents that ultimately answered incorrectly took up to 50% more steps, consumed ~40% more compute, and incurred ~40%
- “CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis” has been accepted as a poster at KDD 2026 (@kdd_news) — one of the leading conferences for AI, machine learning, and data mining. At KDD? Meet with Sentient researcher Darshan Tank (@TankDarshan7) to learn why
- Most AI agents don't just get the math wrong. They get the number from the wrong place: outdated table, mismatched row, wrong year. That's why @abraxasnz13 built an agent that won't calculate until it knows where the number came from ↓
- You can't benchmark your way around a hard problem. Across 13,000+ agent runs in the OfficeQA Public Challenge, Sentient researchers @iamnamanvats and Deep Halder found that different harnesses agreed 88-93% of the time on which tasks succeeded and which failed. TLDR:

