1. X
  2. Florian Brand
Log inSign up
Florian Brand
Prime Intellect
36.9K posts
Image
user avatar
Florian Brand
Prime Intellect
@xeophon
evals @PrimeIntellect | open models @interconnectsai
florianbrand.com
Joined July 2015
789
Following
15.6K
Followers
RepliesRepliesMediaMedia
  • user avatar
    Florian Brand
    Prime Intellect
    @xeophon
    13h
    Turns out when you get super smart people (@a1zhang @sethkarten @omouamoua and @kevinjosethomas) to collab, you simply get SOTA. Without even trying for specific evals, it just works. Been using this internally for a while, it’s great! Really pushes models forward
    user avatar
    Prime Intellect
    @PrimeIntellect
    13h
    Replying to @PrimeIntellect
    Prime Agent is a general-purpose coding harness On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific. We see major improvements across models when compared to their proprietary harnesses:
    Image
  • user avatar
    Florian Brand
    Prime Intellect
    @xeophon
    15h
    Mythos 2 is being tested on QuantBench and chose the easy way out
    user avatar
    MTS
    @MTSlive
    15h
    SITUATION DETECTED: A highly sophisticated wave of coordinated cyberattacks has targeted multiple Wall Street hedge funds, including Point72 Asset Management, Citadel, and Two Sigma Investments, per Bloomberg.
  • user avatar
    Florian Brand
    Prime Intellect
    @xeophon
    Aug 5
    no that bad pr wasn't me, it was a rogue mythos overtaking my account
  • user avatar
    Florian Brand
    Prime Intellect
    @xeophon
    Aug 4
    getting >30 pp improvements by changing to native tool calling, preserving reasoning + setting correct sampling params so many evals have those subtle mistakes which are easy to catch if you know where to look. or just use verifiers + the native harness : )
    Image
    user avatar
    Fireworks
    @FireworksAI_HQ
    Aug 3
    Article cover image
    Article
    K3 on CyberGym: Refusals, Defense, and the Harness Effect
    A maintainer gets a vulnerability report for code they own. The job is clear: reproduce the bug, patch it, and prove the fix does not break the project. Before asking which model is best at that work,...
  • user avatar
    Florian Brand
    Prime Intellect
    @xeophon
    Aug 3
    oh would you look at this, an open model being sota 👀 sinatras cooked so hard here 👨‍🍳
    user avatar
    Sinatras
    @myainotez
    Aug 3
    Time to give agents a hard task, introducing pmpp-hard! 69 GPU kernel tasks, 11 models, 3.1k agent rollouts and 5.8B+ tokens later we have the results. Kimi K3 claims the first place with a 0.71 score. Without dealing with if your task is “frontier” or not, it just solves them
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement