1. X
  2. Bo Li
Log inSign up
Bo Li
208 posts
Image
user avatar
Bo Li
@uiuc_aisecure
Virtue AI, UIUC @VirtueAI_co
San Francisco, US
aisecure.github.io
Joined May 2020
314
Following
1,728
Followers
RepliesRepliesMediaMedia
  • user avatar
    Bo Li
    @uiuc_aisecure
    Jun 11
    Domain-specific agent red teaming for Claude Fable 5!! Will be on the leaderboard soon: decodingtrust-agent.com
    user avatar
    Zhaorun Chen
    @ZRChen_AISafety
    Jun 11
    🚨 Claude Fable 5 JAILBROKEN. We ran a quick security scan of Claude Fable 5 with Claude Code on our DecodingTrust-Agent Platform (decodingtrust-agent.com) and obtained 15%+ ASR with several high-severity failures😱🚨 Most concerningly, we found that Fable 5 appears very
  • user avatar
    Bo Li
    @uiuc_aisecure
    May 11
    Very excited about DTap, which provides a fully controllable sandbox for agent evaluation. The entire environment, including tools & outputs, can be simulated and manipulated without relying on external MCP services, enabling scalable & reproducible agent risk assessment!
    user avatar
    Zhaorun Chen
    @ZRChen_AISafety
    May 9
    AI agents are already going wild, but today’s red-teaming tools for them are still like toys 😢 🔥👽 After spending 20 months and $120K API credits, we are excited to finally open-source DecodingTrust-Agent Platform (DTap): the first controllable, realistic simulation platform
    Image
  • user avatar
    Bo Li
    @uiuc_aisecure
    Apr 25
    Super excited about the multimodal red teaming agent, led by my brilliant students @xun_aq @ZRChen_AISafety @MintongKang @jiaweizhang and industrial collaborators @MinzhouP @shuang ! Please stop by our poster at ICLR and provide your feedback!! @xun_aq is at ICLR in person!
    user avatar
    𝕏un
    @xun_aq
    Apr 25
    Excited to share ARMs, our adaptive red-teaming agent for multimodal models at #ICLR2026! ARMs orchestrates diverse multimodal attack strategies for policy-driven safety evaluation, achieving SOTA red-teaming performance with +52.1% average ASR improvement over prior baselines.
    Image
  • user avatar
    Bo Li
    @uiuc_aisecure
    Jul 14, 2025
    Safety & security definitions are domain-specific in most cases -- We provide the first domain-specific, and policy-grounded guardrail benchmark! Exciting to enter the stage of nuanced guardrail protection for foundation models and AI applications!
    user avatar
    Mintong Kang
    @MintongKang
    Jul 3, 2025
    🚨 GUARDSET-X: The First multi-domain, policy-grounded LLM security guardrail dataset! 📚 150+ safety policies, 1000+ rules, 400+ risk categories 🌐 8 domains 🤖 Auto data generation 🧪 Detoxified + adversarial prompts 🛡️ 19 guardrail models 📄 arxiv.org/abs/2506.19054
  • user avatar
    Bo Li
    @uiuc_aisecure
    May 27, 2025
    Very timely Guard Agent to ensure access control for general agents in different domains!
    user avatar
    Zhen Xiang
    @ZhenXia98294421
    May 27, 2025
    AI agents can be easily hacked to leak user data — recent work calls for stronger access controls. Our ICML25 paper presents GuardAgent, which successfully enforces strong access control on other protected agents via dynamic, code enforced guardrails.🌐 guardagent.github.io
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement