Pinned
Joined March 2017
- yoonholee.com/blog/2026/tast… A few months ago, I wrote this post after a conversation with @Kangwook_Lee . I was mostly complaining that "taste" (research or otherwise) seems to point to something really important, but we have no good way to measure or train for it. This is
- Agreed with everything Omar said; a few other thoughts: - I think RSI will likely start in the harness layer rather than pre/post-training code: it's the cheapest place for an agent to try out hypotheses and fork + modify its own behavior at scale. See e.g. DGM @jennyzhangzt -The answer is deceptively simple imo. There’s a notion of a general-purpose harness (like, say, CoT reasoning, proper compaction, or an RLM) that’s essentially just what *the* model should become. A better “model” in the API will eventually wrap it and train accordingly. And a
- Meta-Harness is a spotlight at the AIWILD workshop @ ICML! I'm not attending this year, but @Kangwook_Lee will be there to present the poster: Sat 15:40-17:00, Hall A, poster #108.How can we autonomously improve LLM harnesses on problems humans are actively working on? Doing so requires solving a hard, long-horizon credit-assignment problem over all prior code, traces, and scores. Announcing Meta-Harness: a method for optimizing harnesses end-to-end
- Thanks for discussing meta-harness @lilianweng ! A very clearly written summary of the state of harness engineering + RSI. I often get the question"what can a harness actually improve?" which i only had partial answers to; this post lays out the whole picture in one placenew post on harness engineering for AI self-improvement: lilianweng.github.io/posts/2026-07-… It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter




