1. X
  2. Anya Sims
Log inSign up
Anya Sims
20 posts
user avatar

Anya Sims

@anyaasims
PhD student @UniofOxford+@FLAIR_Ox supervised by @yeewhye and @j_foerst. Prev interned @graphcoreai; placement @CambridgeMLG. Deep learning, LLMs x RL, meta-RL
Oxford, England
anyasims.github.io
Joined September 2022
151
Following
172
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Anya Sims
    @anyaasims
    Aug 14
    Excited to share our new work from Inherent where we train a 27B AI Scientist model to replicate research papers, and get it to outperform Claude Opus 4.8 and GPT-5.5 Codex! Lots of detail on stabilizing RL in complex, under-specified, long-horizon tasks! arxiv.org/abs/2608.13331
    user avatar
    Inherent
    @inherent_labs
    Aug 14
    1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
    Image
  • user avatar
    Anya Sims
    @anyaasims
    Aug 2, 2025
    Introducing GEM💎: A version of OpenAI Gym for LLMs with everything you need to test out new ideas in multi-turn RL×LLMs!🔥 (Diverse suite of environments, tool integration, async, multi-env training, a standardized interface, simple example scripts (Oat & Verl), and more!) See👇
    user avatar
    Zichen Liu
    @zzlccc
    Aug 1, 2025
    In the era of experience, we're training LLM agents with RL — but something's missing... We miss the good old Gym! So we built 💎GEM: a suite of environments for training LLM 𝚐𝚎𝚗𝚎𝚛𝚊𝚕𝚒𝚜𝚝𝚜. Let’s build the Gym for LLMs, together: axon-rl.notion.site/gem
    Image
  • user avatar
    Anya Sims
    @anyaasims
    Jun 10, 2025
    🎉Excited to share our new work: "StochasTok: Improving Fine-Grained Subword Understanding in LLMs"!🎉 We rethink tokenization and allow LLMs to ‘see inside’ individual tokens by stochastically splitting them into equivalent pairs. See👇! Thanks to my incredible co-authors!🥰
    user avatar
    Cong Lu
    Recursive
    @cong_ml
    Jun 10, 2025
    🚀Introducing “StochasTok: Improving Fine-Grained Subword Understanding in LLMs”!🚀 LLMs are incredible but still struggle disproportionately with subword tasks, e.g., for character counts, wordplay, multi-digit numbers, fixing typos… Enter StochasTok, led by @anyaasims! [1/]
    Image
  • user avatar
    Anya Sims
    @anyaasims
    Dec 3, 2024
    🎉 Excited to share our paper "The Edge-of-Reach Problem in Offline MBRL" has been accepted to #NeurIPS! 🌟 Looking forward to Vancouver! We reveal why offline MBRL methods work (or fail) and introduce a robust solution: RAVL 🚀 🧵 Let's dive in! [1/N]
  • user avatar
    Anya Sims
    @anyaasims
    Feb 23, 2024
    ‼️Offline MBRL fails with the perfect dynamics model‼️ Check out our new paper where we expose a catastrophic “edge-of-reach” pathology ➡️ and completely re-explain why current offline MBRL works/fails! 🤩 Paper: arxiv.org/abs/2402.12527 Code: github.com/anyasims/edge-… 👇
    user avatar
    Cong Lu
    Recursive
    @cong_ml
    Feb 23, 2024
    🚨 Model-based methods for offline RL aren’t working for the reasons you think! 🚨 In our new work, led by @anyaasims, we uncover a hidden “edge-of-reach” pathology which we show is the actual reason why offline MBRL methods work or fail! Let's dive in! 🧵 [1/N]
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement