Log inSign up
Boyi Wei
172 posts
Boyi Wei profile banner
@boyiwei

Boyi Wei

@boyiwei
PhD student @Princeton @PrincetonCITP.
Princeton, NJ
boyiwei.com
Joined February 2020
856
Following
527
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @boyiwei
    Boyi Wei
    @boyiwei
    Jun 3, 2025
    Are static evaluations enough to reflect the risks of offensive cybersecurity agents? 🤔 We show that the answer is no! 😯 Even with minimal compute, adversaries can significantly boost offensive cybersecurity performance -- without any external assistance! 🧵👇[1/n]
    Image
    GIF
    1
  • @boyiwei
    Boyi Wei
    @boyiwei
    Apr 24
    Great work! I like the idea of learning an option-style controller that decides when to keep or switch expert sets instead of switching experts at nearly every token. Check this out!
    @Zeyu_Shen_yo
    Zeyu Shen
    @Zeyu_Shen_yo
    Apr 24
    New paper! 🧵 Modern MoE LLMs switch their active experts at >94% of token positions. This makes memory offloading more expensive We show you can cheaply convert pretrained MoEs (like gpt-oss-20b) into temporally extended ones, dropping switch rates from >50% to <5% while
    Image
  • @boyiwei
    Boyi Wei
    @boyiwei
    Nov 28, 2025
    I'll be at #NeurIPS2025 from Dec 2 to Dec 7! Happy to catch up and meet new friends, especially those who are interested in Agents (self-improvement, scientific discovery) and AI alignment! I will also present two papers:
    1
  • @boyiwei
    Boyi Wei
    @boyiwei
    Nov 21, 2025
    For Bio-foundation models, data filtering cannot completely prevent the model from being misused in some cases. We show that, even for the models trained with data filtering, we can still be able to recover harmful capabilities via probing or fine-tuning. Check this out!
    @UdariMadhu
    Udari Madhushani Sehwag
    @UdariMadhu
    Nov 21, 2025
    Our new research from @scale_AI reveals that harmful biological knowledge can persist inside bio-foundation models even after filtering. We introduce BioRiskEval, the first comprehensive framework built to assess dual-use risk in these models using a realistic adversarial threat
    Image
  • @boyiwei
    Boyi Wei
    @boyiwei
    Oct 7, 2025
    "Misevolve" can happen unintentionally when agents improve themselves, leading to undesired (mostly harmful) outcomes. As people focus more on self-improving agents, alignment strategies should also be adaptive to avoid emergent risks. Check this out if you are interested!
    @jsonren00
    Qihan Ren
    @jsonren00
    Oct 7, 2025
    [1/8] New risk in self-evolving agents: "Misevolution"—when self-evolution unintentionally deviates and causes harm. We found this in various evolutionary paths (model, memory, tool, workflow), even with SOTA LLMs. E.g. a coding agent's ASR surged from 0.6% to 20.6% after memory
    Image
    Image
    1
Advertisement
Advertisement