Log inSign up
Longxu Dou
126 posts
@LongxuDou

Longxu Dou

@LongxuDou
Researcher @TencentHunyuan | ex-@SeaAIL
longxudou.github.io
Joined February 2016
701
Following
269
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @LongxuDou
    Longxu Dou
    @LongxuDou
    Feb 19, 2025
    🚀 Excited to share our technical report on the Southeast Asian multilingual model Sailor2 and its latest updates! Our 49-page report details Sailor2's development journey, including multilingual data cleaning, small model data mixture simulations, multi-stage continual
    Image
    Image
    Image
    Image
    9
  • @LongxuDou
    Longxu Dou
    @LongxuDou
    Jul 14
    Grok 4.20 ranked last at 0.080. With Grok 4.5, it jumped to 0.505 — 13/46 tasks, from last to first on the long-duration agent track. Huge progress from the xAI team. Still plenty of headroom: 29 of 46 tasks remain unsolved by any model.
    @Yucheng__Shi
    Yucheng Shi
    @Yucheng__Shi
    Jul 13
    Most agent benchmarks end in minutes. Real terminal work doesn’t. We’re introducing Long-Horizon Terminal-Bench (LHTB), a benchmark designed to measure whether AI agents can sustain progress across hundreds of dependent actions, not just start a task and solve the easy first
    Image
    00:00
  • @LongxuDou
    Longxu Dou
    @LongxuDou
    Jul 12
    Image
    @GoSailGlobal
    Jason Zhu
    00:00
    @GoSailGlobal
    Jason Zhu
  • @LongxuDou
    Longxu Dou
    @LongxuDou
    Dec 17, 2025
    🚀We propose Reptile, a Terminal Agent🤖️that enables interaction with an LLM agent directly in your terminal. The agent can execute any command or custom CLI tool to accomplish tasks, and users can define their own tools and commands for the agent to utilize. ✨What Makes
    Image
    4
  • @LongxuDou
    Longxu Dou
    @LongxuDou
    Aug 9, 2025
    🚀 Diffusion Language Models are super Data Learners!
    @NiJinjie
    Jinjie Ni
    @NiJinjie
    Aug 9, 2025
    Token crisis: solved. ✅ We pre-trained diffusion language models (DLMs) vs. autoregressive (AR) models from scratch — up to 8B params, 480B tokens, 480 epochs. Findings: > DLMs beat AR when tokens are limited, with >3× data potential. > A 1B DLM trained on just 1B tokens
    Image
Advertisement
Advertisement