Log inSign up
AlphaSignal
1,328 posts
AlphaSignal profile banner
@AlphaSignalAI

AlphaSignal

@AlphaSignalAI
We help you track, rank, and understand the entire AI industry in real time. Used by 300,000+ developers. 5-min daily AI digest: alphasignal.ai/newsletter
Build your feed
alphasignal.ai
Joined February 2010
340
Following
16.5K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @AlphaSignalAI
    AlphaSignal
    @AlphaSignalAI
    Jul 18
    Kimi K3 is getting called Fable/Sol level, and it's 7th in our tests. Arena Frontend Code: #1 at 1679 points. Artificial Analysis: #3 at Intelligence Index of 57. We ran it the next day on our coding-agent repair harness against GPT-5.6 Sol, Fable 5, Grok 4.5, Opus 4.8,
    Image
    @ArtificialAnlys
    Artificial Analysis
    @ArtificialAnlys
    Jul 16
    Image
    Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open
    142
  • @AlphaSignalAI
    AlphaSignal
    @AlphaSignalAI
    Sep 3
    Your 5-minute cache is already dead when you sit down. API models don't keep the thread. Every turn resends the whole history. A one-token tool result still pays that whole history. On Opus 5, 1 million uncached tokens are $5. A hit is $0.50. Cache-hit discounts are
    Image
    Image
    2
  • @AlphaSignalAI
    AlphaSignal
    @AlphaSignalAI
    Sep 3
    AI agents can write the code, but can they tell when they broke something? We’re hosting Vilhelm von Ehrenheim, Co-founder & Chief AI Officer at @getqatech, for a technical deep dive into the missing half of agentic development: verification. Register → luma.com/moa2f7ru
    Image
  • @AlphaSignalAI
    AlphaSignal
    @AlphaSignalAI
    Sep 2
    Agent frameworks keep changing because tool calling never solved the runtime. A tool schema tells the model how to call one function. It does not decide who owns the loop, where state lives, what counts as done, or how recovery works. @NVIDIAAI's NOOA takes a different
    Image
    3
  • @AlphaSignalAI
    AlphaSignal
    @AlphaSignalAI
    Sep 1
    A moderation model can read your policy, react to it, and still enforce it incorrectly. Tested @MistralAI's 3B Shieldstral on 124 policy decisions. The model clearly responds to runtime rules. Every unfamiliar-policy match scored above its unrelated control. But the failures
    Image
    2
Advertisement
Advertisement