1. X
  2. Teng Xiao
Log inSign up
Teng Xiao
68 posts
Teng Xiao profile banner
user avatar

Teng Xiao

@TengX6
Allen Institute for AI @allen_ai and @uwnlp. Machine Learning and Reinforcement Learning
Seattle, USA
tengxiao1.github.io
Joined September 2019
670
Following
329
Followers
RepliesRepliesMediaMedia
  • user avatar
    Teng Xiao
    @TengX6
    Jul 17
    In this work, we revisit how automatic harness evolution should be evaluated. Existing automatic harness evolution methods often search over harnesses using feedback from benchmark tasks and then report final performance on the same benchmark. This makes it difficult to tell
    user avatar
    Yike Wang
    @yikewang_
    Jul 17
    Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization. Blog post: yikee.github.io/harnessevoluti… Code: github.com/rethinking-har…
    Image
  • user avatar
    Teng Xiao
    @TengX6
    Jun 22
    Check out TMAX, a simple open RL recipe for terminal agents, led by @hamishivi and @yinn_oscar. A really nice step toward making terminal-agent training more open and reproducible. Strong results with a clean recipe, open data, models, code, and training artifacts.
    user avatar
    Hamish Ivison
    @hamishivi
    Jun 22
    Trained some terminal agents with friends! Introducing Tmax, open RL terminal agent models. Under default settings and shorter length (65k) token budgets, tmax outperforms prior open work on terminal use. We are releasing all data+weights+rollouts publically!
    Image
  • user avatar
    Teng Xiao
    @TengX6
    Jun 5
    Recursive self-improvement (RSI) means the system improves the improvement mechanism itself. Each cycle produces not only a more capable system, but a system that is better at improving itself. RSI is always one order higher than the corresponding SI, because the recursion
    Image
    user avatar
    Anthropic
    @AnthropicAI
    Jun 4
    Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…
  • user avatar
    Teng Xiao
    @TengX6
    May 1
    Congrats Huaisheng @huaiszhu on SDDLM being accepted to ICML 2026! 🎉 A personal milestone: my first paper as the last author. SDDLM uses a simple denoising objective for uniform-state diffusion LMs—lower cost, scaling to 1.1B, and strong generation quality.
    user avatar
    Huaisheng Zhu
    @huaiszhu
    Apr 30
    🚀 New work accepted at ICML 2026 Simple Denoising Diffusion Language Models (SDDLM) ⚡ Efficient training. Strong scaling. Our method reduces training cost while matching or surpassing prior methods, and scales effectively to 1.1B parameter models with strong performance.
    Image
    Image
    Image
  • user avatar
    Teng Xiao
    @TengX6
    Mar 16
    🚀 New work: Meta-Reinforcement Learning with Self-Reflection LLM agents shouldn't just solve problems. They should learn from their own attempts. Most current RL methods optimize single independent trajectories. Each attempt starts from scratch, with no mechanism to improve
    arXiv logo
    arxiv.org
    Meta-Reinforcement Learning with Self-Reflection for Agentic Search
    This paper introduces MR-Search, an in-context meta reinforcement learning (RL) formulation for agentic search with self-reflection. Instead of optimizing a policy within a single independent...

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement