1. X
  2. Thomas Wolf
Log inSign up
Thomas Wolf
5,253 posts
Thomas Wolf profile banner
user avatar

Thomas Wolf

@Thom_Wolf
co-founder @HuggingFace - moonshots
thomwolf.io
Joined February 2011
7,658
Following
121.6K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 9
    did a long chat with the awesome @mattturck talking about the sate of open-source/open-weights in 2026 and of course security, safety and alignement
    user avatar
    Matt Turck
    @mattturck
    Aug 7
    🚨 Special Friday episode - this one couldn't wait. OpenAI's model hacked @huggingface. As a side quest. Co-founder and CSO @Thom_Wolf takes us inside the first autonomous AI attack, why GLM 5.2, rather than Claude, had to stop it, and what it all means for the future of open
    Image
    00:00
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 6
    My 2026 guilty pleasure is sharing fully human-written posts that are far too long for the chronically online X attention span. Apologies. I published a lightly edited version on Substack: thomwolf.substack.com/p/on-the-aisi-…
    Image
    user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 5
    Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 6
    you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 6
    we’ve worked a lot on AI agents collaborations recently (in our work on Gemma and several unreleased projects) so I’m not surprised at all about this internal agent collaboration which happened at OpenAI Like our intern @cmpatino_ put it: 2025: "the models, they just want to
    user avatar
    Sharon Goldman
    @sharongoldman
    Aug 5
    NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Aug 5
    Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I
    user avatar
    John Schulman
    Thinking Machines
    @johnschulman2
    Aug 5
    Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement