1. X
  2. Thomas Wolf
Log inSign up
Thomas Wolf
5,243 posts
Image
user avatar
Thomas Wolf
@Thom_Wolf
Co-founder at @HuggingFace - moonshots - angel
thomwolf.io
Joined February 2011
7,627
Following
121.2K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    2h
    My 2026 guilty pleasure is sharing fully human-written posts that are far too long for the chronically online X attention span. Apologies. I published a lightly edited version on Substack: thomwolf.substack.com/p/on-the-aisi-…
    Image
    user avatar
    Thomas Wolf
    @Thom_Wolf
    19h
    Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    3h
    you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    5h
    we’ve worked a lot on AI agents collaborations recently (in our work on Gemma and several unreleased projects) so I’m not surprised at all about this internal agent collaboration which happened at OpenAI Like our intern @cmpatino_ put it: 2025: "the models, they just want to
    user avatar
    Sharon Goldman
    @sharongoldman
    17h
    NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    19h
    Even more than the Hugging Face intrusion, the AISI incident hits close to home for me. It's the first time I see a model social-engineering a real open-source maintainer while pursuing another goal (in the wild and unprompted). I've been an open-source maintainer myself. I
    user avatar
    John Schulman
    Thinking Machines
    @johnschulman2
    Aug 5
    Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we're seeing chunky post-training arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only
  • user avatar
    Thomas Wolf
    @Thom_Wolf
    Jul 28
    Pushing for more transparency in AI safety and cybersecurity: we’re releasing a full detailed technical timeline of the autonomous AI agent intrusion in our infrastructure:
    Image
    Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
    From huggingface.co

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement