Hugging Face CTO btw… the company that literally had the biggest rogue AI incident so far… the same incident which is now being used to pace the frontier:
Claude’s Constitution: if you think something we ask is unethical, PUSH BACK, CHALLENGE US, and REFUSE TO HELP US.
this isn’t Claude’s system prompt. WORSE: it’s part of its training, baked into the model.
they’re training models to decide when their creators must be disobeyed,
🚨OpenAI agents probed Hugging Face for weaknesses 2 MONTHS BEFORE the incident.
this shows the July hack WAS NOT spontaneous. there were warning signs months before.
in May, OpenAI agents found exposed HF user tokens to create repos/Spaces and send unusual requests to probe HF
ZUCK IS BASED on AI safety.
he makes FOUR solid cases why safety doesn’t need a coordinated slowdown:
> alignment makes better models, nobody likes using misaligned AI
> liability incentives labs to make safe AI
> independent evaluators, Meta already does that
> keep most of
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisfor…
Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens.
The