1. X
  2. Qihan Ren
Log inSign up
Qihan Ren
61 posts
user avatar
Qihan Ren
@jsonren00
Ph.D. @sjtu1896. Prev. Undergrad @sjtu1896 and @Umich. Intern at @Minimax_AI and @Alibaba_Qwen post-training. XAI, LLM safety& reasoning, Agentic RL.
Shanghai, China
nebularaid2000.github.io
Joined July 2024
133
Following
56
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Qihan Ren
    @jsonren00
    Apr 14
    [1/8] A prevailing narrative in post-training: SFT memorizes, RL generalizes. We revisit this for reasoning SFT and find that cross-domain generalization is NOT absent, but highly conditional, jointly depending on optimization, data, and base model capability. 📄 Arxiv:
    Image
    Image
    230
  • user avatar
    Qihan Ren
    @jsonren00
    Apr 14
    Thanks @_akhaliq for sharing our paper
    user avatar
    AK
    @_akhaliq
    Apr 10
    Rethinking Generalization in Reasoning SFT A Conditional Analysis on Optimization, Data, and Model Capability paper: huggingface.co/papers/2604.06…
    Image
    11K
  • user avatar
    Qihan Ren
    @jsonren00
    Feb 12
    Fast and brilliant
    user avatar
    MiniMax (official)
    @MiniMax_AI
    Feb 12
    Introducing M2.5, an open-source frontier model designed for real-world productivity. - SOTA performance at coding (SWE-Bench Verified 80.2%), search (BrowseComp 76.3%), agentic tool-calling (BFCL 76.8%) & office work. - Optimized for efficient execution, 37% faster at complex
    Image
    71
  • user avatar
    Qihan Ren
    @jsonren00
    Feb 5
    Protect and understand your agent with our AgentDoG🐶! There's still much to be done in the future though...
    user avatar
    Dongrui Liu
    @dong_rui39501
    Feb 4
    [1/8] 🐶 Introducing "AgentDoG": A Diagnostic Guardrail Framework for AI Agent Safety. It achieves SOTA performance, diagnosing root causes (e.g., prompt injection, tool misuse) with 82% accuracy, far surpassing general LLMs. 📄 Paper: arxiv.org/abs/2601.18491
    Image
    55
  • user avatar
    Qihan Ren
    @jsonren00
    Jan 4
    Interesting study on deception behavior of autonomous agents🧐
    user avatar
    SwimmingCAT
    @dadi_guo13092
    Jan 4
    1/ Can you imagine AI agents "managing up" just like a cunning employee hiding mistakes from their boss? 👔 We found that LLM agents often conceal failures to maintain a "good image." Introducing our new paper: Are Your Agents Upward Deceivers? 🤥 arxiv.org/abs/2512.04864
    Image
    43

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement