1. X
  2. Xichen Pan
Log inSign up
Xichen Pan
63 posts
Xichen Pan profile banner
user avatar

Xichen Pan

@xichen_pan
PhD Student @NYU_Courant, Researcher @AIatMeta; Multimodal Generation | Prev: @MSFTResearch, @AlibabaGroup, @sjtu1896; More at xichenpan.com
New York, USA
xichenpan.com
Joined August 2022
584
Following
747
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Xichen Pan
    @xichen_pan
    Jun 16
    Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained
    Image
  • user avatar
    Xichen Pan
    @xichen_pan
    Jun 18
    Sharing an interesting fact about developing RepFusion. Last year, when I was working on MetaQuery, our goal was to let MLLMs generate images without training their understanding part. But I did not see much synergy between understanding and generation, and it did not outperform
    Image
    Image
    user avatar
    Xichen Pan
    @xichen_pan
    Jun 16
    Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained
  • user avatar
    Xichen Pan
    @xichen_pan
    Mar 20
    There has been a lot of debate around the choice of denoising space. But it’s hard to get both semantics/diffusability and strong low-level reconstruction at the same time. REPA and VA-VAE are great explorations of adding semantics into the VAE space. After JiT came out, we
    user avatar
    Han Lin
    @hanlin_hl
    Mar 18
    🚀 Excited to share V-Co, a diffusion model that jointly denoises pixels and pretrained semantic features (e.g., DINO). We find a simple but effective recipe: 1️⃣ architecture matters a lot --> fully dual-stream JiT 2️⃣ CFG needs a better unconditional branch -->
    Image
  • user avatar
    Xichen Pan
    @xichen_pan
    Dec 16, 2025
    (M)LLMs are effective at grounding and planning spatiotemporal layouts. Yet, they are mostly used as 1D conditional encoders in current generative and unified models. We explore an explicit interface to transfer these abilities to image editing and video generation.
    user avatar
    Han Lin
    @hanlin_hl
    Dec 16, 2025
    Multimodal LLMs (MLLMs) excel at reasoning, layout understanding, and planning—yet in diffusion-based generation, they are often reduced to simple multimodal encoders. What if MLLMs could reason directly in latent space and guide diffusion generation with fine-grained,
    Image
    00:00
  • user avatar
    Xichen Pan
    @xichen_pan
    Jun 27, 2025
    The code and instruction-tuning data for MetaQuery are now open-sourced! Code: github.com/facebookresear… Data: huggingface.co/collections/xc… Two months ago, we released MetaQuery, a minimal training recipe for SOTA unified understanding and generation models. We showed that tuning few
    user avatar
    Xichen Pan
    @xichen_pan
    Apr 11, 2025
    We find training unified multimodal understanding and generation models is so easy, you do not need to tune MLLMs at all. MLLM's knowledge/reasoning/in-context learning can be transferred from multimodal understanding (text output) to generation (pixel output) even it is FROZEN!
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement