1. X
  2. Robert Kirk
Log inSign up
Robert Kirk
554 posts
user avatar

Robert Kirk

@_robertkirk
Alignment Red Team at @AISecurityInst. Prev. PhD Student @ucl_dark
robertkirk.github.io
Joined January 2020
345
Following
2,037
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Robert Kirk
    @_robertkirk
    Apr 27
    We evaluated Claude Mythos Preview, Opus 4.7 and other models with our updated alignment evaluation methodology, including a new continuation eval, improved evaluation and prefill awareness measurements. Details including new methodology in 🧵:
    user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Apr 27
    As part of our work on assessing AI loss-of-control risks, we collaborated with @AnthropicAI to pilot alignment evals on models including pre-release snapshots of Mythos Preview and Opus 4.7. We ask: could an AI agent used inside a frontier lab sabotage safety research? 🧵
    Image
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 29
    Come work with us! It's the best place I've worked by far, the team are exceptional and lots of fun, and @abbydcruz__ is incredible.
    user avatar
    Abby D'Cruz
    @abbydcruz__
    Jul 28
    The Red Team at @AISecurityInst is hiring! We test the misuse safeguards, control measures, and alignment of frontier models. We're looking for a high-agency generalist to run the delivery function that underpins our impact 🧵
    Image
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 24
    We tested Opus 5 for whether it would sabotage safety research @AISecurityInst We saw no unprompted sabotage and very low rates of continuing sabotage. However, it’s the best model we’ve tested at distinguishing our evals from deployment when prompted. 🧵 on these results
    Image
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 21
    Very cool and impressive work from @apolloresearch and @OpenAI – excited to see it out publicly!
    user avatar
    Apollo Research
    @ApolloResearch
    Jul 21
    Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards? Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
    Image
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 9
    @AISecurityInst did our first pre-deployment alignment testing with OpenAI for GPT 5.6 Sol! 3 key takeaways 🧵
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement