1. X
  2. Thomas Read
Log inSign up
Thomas Read
126 posts
user avatar
Thomas Read
@thjread
Interpretability researcher at UK AISI
London, UK
blog.thjread.com
Joined November 2018
227
Following
192
Followers
RepliesRepliesMediaMedia
  • user avatar
    Thomas Read
    @thjread
    May 21
    I helped write this report on oversight of AI systems and how it could degrade - it's a great overview, and a good guide to what research directions might help us maintain the level of oversight we enjoy today
    user avatar
    AI Security Institute (AISI)
    @AISecurityInst
    May 21
    The safety of advanced AI systems increasingly depends on the ability to oversee them. Our new report examines today’s AI oversight landscape, finding many pathways likely to lead to its degradation.🧵
    Image
  • user avatar
    Thomas Read
    @thjread
    Apr 10
    New from the UK AISI Model Transparency team: we replicated Anthropic's steering approach for suppressing evaluation awareness. Our most surprising finding: "control" steering vectors (about books on shelves!) can have effects as large as deliberately designed ones. 🧵
    Image
  • user avatar
    Thomas Read
    @thjread
    Mar 30
    more exciting work from my team at UK AISI! an open source reproduction of "Natural Emergent Misalignment from Reward Hacking in Production RL"
    user avatar
    7vik
    @satvikgolechha
    Mar 30
    Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness. 🧵 (1 of 7)
    Image
  • user avatar
    Thomas Read
    @thjread
    Dec 10, 2025
    First paper from my team at UK AISI! Excited to have this out there - we have some really great model organisms of conditional underperformance, and tried a lot of different detection techniques to see what works in an adversarial setting
    user avatar
    Jordan Taylor
    @JordanTensor
    Dec 9, 2025
    NEW PAPER from UK AISI Model Transparency team: Could we catch AI models that hide their capabilities? We ran an auditing game to find out. The red team built sandbagging models. The blue team tried to catch them. The red team won. Why? 🧵1/17
    Image
  • user avatar
    Thomas Read
    @thjread
    Jul 10, 2025
    I've recently started working at UK AISI! check out what my team was working on before I arrived, looking forward to bringing you updates on this work in the future!
    user avatar
    Joseph Bloom
    @JBloomAus
    Jul 10, 2025
    🧵 1/13 My new team at UK AISI - the White Box Control Team - has released progress updates! We've been investigating whether AI systems could deliberately underperform on evaluations without us noticing. Key findings below 👇

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement