1. X
  2. Eric Wong
Log inSign up
Eric Wong
194 posts
user avatar
Eric Wong
@RICEric22
Assistant professor at University of Pennsylvania. Machine learning, optimization, robustness & interpretability. profericwong.bsky.social
Philadelphia, PA
cis.upenn.edu/~exwong/
Joined July 2009
110
Following
1,851
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Eric Wong
    @RICEric22
    Apr 10
    You may have heard agentic benchmarks can be gamed. Turns out, leading solutions already do this. With Meerkat, our framework for large-scale trace auditing, we found thousands of clear cheating instances, likely from unsupervised vibe coding. debugml.github.io/cheating-agent…
    user avatar
    Adam Stein
    @adamlsteinl
    Apr 10
    We found widespread cheating on popular agent benchmarks, affecting 28+ submissions across 9 benchmarks and thousands of agent runs. Surprisingly, the top 3 submissions on Terminal-Bench 2 are all cheating! Here's what we found 🧵
    Image
    1.3K
  • user avatar
    Eric Wong
    @RICEric22
    May 11
    Recent work led by @WeiqiuYou at the intersection of VLM reasoning and the surgical problem of identifying the critical-view-of-safety! A fun (and eye-opening) collaboration with @Laparoscopes for explainable AI in surgery problems.
    user avatar
    PennSurgery
    @pennsurgery
    May 11
    New collaborative work b/w @GRASPlab @PCASOLab @Laparoscopes &@CIS_Penn @RICEric22 shows how #XAI methods for structured reasoning improves large VLM performance in identifying CVS anatomy in lap chole Full text at rdcu.be/fh1M7
    Image
    744
  • user avatar
    Eric Wong
    @RICEric22
    Jul 10, 2025
    LLM ignoring instructions? Make it listen with InstABoost. ✅ Simple: Steer your model in 5 lines of code ✅ Effective: Outperforms latent steering & prompt-only methods ✅ Grounded: Based on our mechanistic theory on rule-following (LogicBreaks) Blog: debugml.github.io/instaboost
    user avatar
    Adam Stein
    @adamlsteinl
    Jul 10, 2025
    Excited to share our new paper: "Instruction Following by Boosting Attention of Large Language Models"! We introduce Instruction Attention Boosting (InstABoost), a simple yet powerful method to steer LLM behavior by making them pay more attention to instructions. (🧵1/7)
    Image
    GIF
    1.9K
  • user avatar
    Eric Wong
    @RICEric22
    Jul 9, 2024
    Why can safety rules in LLMs be jailbroken? In LogicBreaks, we study the fundamental mechanism behind rule subversion in LLMs. Our theory explains how one can force LLMs to suppress rules/knowledge and infer absurd facts--and it mirrors real jailbreaks! debugml.github.io/logicbreaks/
    Image
    user avatar
    Anton Xue
    @AntonXue
    Jul 9, 2024
    Happy to present our recent work: "Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference" 📝 Blog: debugml.github.io/logicbreaks/ 🧐 arXiv: arxiv.org/abs/2407.00075 🤖 Code: github.com/AntonXue/tf_lo…
    4.1K
  • user avatar
    Eric Wong
    @RICEric22
    Jul 5, 2024
    Traditional concept vectors used to explain deep representations fail to compose when combined, i.e. 🐤(small) +🦢(white) =🦩(big & colorful)❌ We propose CCE: a method for extracting *composable* concepts, i.e. 🐤(small) +🦢(white) =🕊️(small & white)✅ debugml.github.io/compositional-…
    user avatar
    Adam Stein
    @adamlsteinl
    Jul 5, 2024
    Excited to present our ICML 2024 paper: "Towards Compositionality in Concept Learning"! 🔗 Blog: debugml.github.io/compositional-… 📄 Paper: arxiv.org/abs/2406.18534 💻 Code: github.com/adaminsky/comp…
    Image
    5.1K
  • See @RICEric22's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement