PartConcepts

A Unified Mechanism for Fine-Grained Part Localization and Generation

NeurIPS 2026 🎉

also presented at GCV workshop, CVPR 2026

Vaibhav Agrawal1, Varghese P Kuruvilla1, Harsh Rangwani2, Ravi Kiran S1

1 IIIT Hyderabad    2 Adobe Research, Bengaluru

Code and arXiv links coming soon

Motivation

Diffusion latents have strong semantics at the object-level, e.g., dog.

A dog standing in grass. Cross-attention map for the dog token.
Cross-attention map with the dog token.

But this fails to extend to the part-level granularity, e.g., left front leg.

A dog standing in grass. Cross-attention map for the left token. Cross-attention map for the front token. Cross-attention map for the leg token.
Cross-attention maps with the left, front, and leg tokens.

The poor part-level semantics correlate with several capability gaps in existing T2I diffusion models.

Capability gaps in existing text-to-image diffusion models: poor concept segmentation and poor part-level prompt following.

Helbling, Alec, et al. "Conceptattention: Diffusion transformers learn highly interpretable features." ICML 2025.

Key motivating observations

Prior work has shown that diffusion latents enable dense semantic correspondenceSemantic correspondence points on a grey cat.Semantic correspondence points transferred to a second cat.Semantic correspondence matches the same visual parts across different images and poses using internal diffusion features. Figure adapted from Tang et al., Emergent Correspondence from Image Diffusion, NeurIPS 2023; see also Zhang et al., Diffusion Features as Semantic Representations, ICML 2023.; hence possess the necessary resolution for part-level discrimination.

The dispersed cross-attention mapsA dog standing in grass.Attention map for the left token.leftAttention map for the front token.frontAttention map for the leg token.legCross-attention is very dispersed for the part phrase left front leg. indicate that the text-image interaction is a bottleneck, failing to leverage the dense diffusion latents.

Research question

Our observations suggest that part phrases (e.g., left ear, right eye) are not well represented on the text-image manifold. Our core research question, therefore, is the following:

Can we introduce textual part concepts into the T2I model, while preserving its base capabilities?

Method

Given an input text prompt with part phrases (e.g., ‘left ear’), our goal is to make those phrases attend strongly to the correct part regions in the image — enabling both precise part-level localization and fine-grained instruction following in T2I generation.

Encoding textual part descriptions into PartConcepts

  • Each part-phrase (e.g., left ear) is mean pooled into a compact PartConcept token, which represents a coherent semantic signal rather than disparate tokens.
  • Selective layers of the text encoder are LoRA finetuned; this enables the PartConcept token to learn part-level semantics.
PartConcept representation formed by mean pooling the textual embeddings of a part phrase.

The training objectives

Training pipeline: noisy visual latents and text with PartConcept tokens pass through mmDiT, attention maps are decoded to segmentation masks, and flow loss preserves the generative prior.
  1. Noisy visual latents, and text + PartConcept tokens are input to the mmDiT.

  2. Layer-wise cross attention maps are extracted.

  3. Layer-wise concatenated attention maps are decoded to segmentation masks, and penalized using Dice loss.

  4. Flow loss on the velocity prediction; this preserves the generative capabilities.

Balancing localization and generation

  • Both the dice loss (localization) and flow loss (generation) are balanced during training:

    L = Ldice + λLflow
  • λ governs a tradeoff between localization accuracy and generation quality, as shown in the ablation below.

Ablation study showing the effect of the flow loss weight on localization accuracy and generation quality.

Evaluating PartConcepts

  • We evaluate our method PartConcepts for localization and generation capabilities
    • Part instance segmentation
    • Part-level instruction following in text-to-image generation

Part-level instruction following

Generating images that follow fine-grained instructions about specific object parts.

Examples of part-level instruction following with PartConcepts.

Part instance segmentation

Localizing and segmenting individual object parts from open-vocabulary text prompts.

Examples of part instance segmentation with PartConcepts.

Summary of experiments

  • PartConcepts enables state-of-the-art part instance segmentation
    • respects spatial organization of part instances (left front leg vs. right back leg)
  • The textual attributes in the prompt naturally bind to the PartConcept token
    • e.g., purple binds to left lower arm
    • without need for explicit architectural mechanisms for attribute binding
  • The PartConcept token behaves similar to a standard text embedding, but enriched with part-level semantics!
Attribute binding example showing PartConcept attention for right upper arm and left lower arm.

BibTeX

@inproceedings{agarwal2026partconcepts,
  title     = {{\textsc{PartConcepts}}: A Unified Mechanism for Fine-Grained Part
               Localization and Generation},
  author    = {Agrawal, Vaibhav and Kuruvilla, Varghese P and Rangwani, Harsh
               and Sarvadevabhatla, Ravi Kiran},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  note      = {To appear}
}