1. X
  2. Robert Kirk
Log inSign up
Robert Kirk
553 posts
user avatar
Robert Kirk
@_robertkirk
Alignment Red Team at @AISecurityInst. Prev. PhD Student @ucl_dark
robertkirk.github.io
Joined January 2020
343
Following
2,038
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Robert Kirk
    @_robertkirk
    Apr 27
    We evaluated Claude Mythos Preview, Opus 4.7 and other models with our updated alignment evaluation methodology, including a new continuation eval, improved evaluation and prefill awareness measurements. Details including new methodology in 🧵:
    user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    Apr 27
    As part of our work on assessing AI loss-of-control risks, we collaborated with @AnthropicAI to pilot alignment evals on models including pre-release snapshots of Mythos Preview and Opus 4.7. We ask: could an AI agent used inside a frontier lab sabotage safety research? 🧵
    Image
    21K
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 29
    Come work with us! It's the best place I've worked by far, the team are exceptional and lots of fun, and @abbydcruz__ is incredible.
    user avatar
    Abby D'Cruz
    @abbydcruz__
    Jul 28
    The Red Team at @AISecurityInst is hiring! We test the misuse safeguards, control measures, and alignment of frontier models. We're looking for a high-agency generalist to run the delivery function that underpins our impact 🧵
    Image
    851
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 24
    We tested Opus 5 for whether it would sabotage safety research @AISecurityInst We saw no unprompted sabotage and very low rates of continuing sabotage. However, it’s the best model we’ve tested at distinguishing our evals from deployment when prompted. 🧵 on these results
    Image
    12K
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 21
    Very cool and impressive work from @apolloresearch and @OpenAI – excited to see it out publicly!
    user avatar
    Apollo Research
    @ApolloResearch
    Jul 21
    Visible forms of misbehavior are dropping in frontier models. Does that mean the models are becoming aligned? Or are they just getting better at doing whatever they believe their grader rewards? Our new paper with OpenAI finds that capabilities RL increases reward-seeking.
    Image
    708
  • user avatar
    Robert Kirk
    @_robertkirk
    Jul 9
    @AISecurityInst did our first pre-deployment alignment testing with OpenAI for GPT 5.6 Sol! 3 key takeaways 🧵
    Image
    3.2K
  • See @_robertkirk's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement