Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation (ECCV 2026)

Project Page | arXiv | Code

NaLA teaser


Introduction

This is the official implementation of NaLA, a 3D-native LLM layout agent for high-quality 3D scene generation. NaLA encodes point clouds of scenes and assets as native 3D tokens and generates placements through a coarse-to-fine pose decoding strategy, enabling end-to-end reasoning over 3D geometry for physically and semantically plausible layouts.

Implementation notes. Default backbone: Qwen/Qwen2.5-7B-Instruct. Default asset encoder: Point-MAE (PointBERT supported). Training mix: 3D-FRONT, Imaginarium, and M3DLayout via --dataset_to_use mixed_all.


Project structure

NaLA-code/
├── assets/                 # Teaser figure
├── checkpoints/encoders/   # Point-MAE / PointBERT / SPFormer weights
├── configs/                # DeepSpeed & model configs
├── docs/                   # Dataset preparation notes
├── scripts/                # train_ds.py, train.py, infer.py
├── shells/                 # Multi-GPU launch scripts
├── tools/                  # Point-cloud & metadata preprocessing
└── src/nala/
    ├── data/               # Dataset & collators
    ├── models/             # LayoutLLM + point encoders
    ├── post_processing/    # Optional physics cleanup
    └── utils/

Environment

Install the package and dependencies with:

git clone https://github.com/adamcwan/NaLA-code.git
cd NaLA-code
pip install -e .

Dependencies: PyTorch (≥2.1), Transformers, PEFT, DeepSpeed, trimesh, spconv (Linux). LLM weights: Qwen/Qwen2.5-7B-Instruct from Hugging Face.

Encoder weights

Download the pretrained encoders below and place them under checkpoints/encoders/ (rename to the local filenames if needed):

File Role Source
pointMAE_pretrain.pth Asset encoder (Point-MAE) Point-MAE (pretrain.pth)
point_bert_v1.2.pt Asset encoder (PointBERT, optional) Hugging Face
spformer_encoder_uniform_superpoints.pth Scene encoder (SPFormer) SPFormer

Datasets

Prepare 3D-FRONT, Imaginarium, and M3DLayout, then convert them to the expected layout. See docs/dataset.md.


Training

Start multi-GPU training with:

export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export NALA_DATA_ROOT=data
bash shells/run_ds_mixed_all.sh

Trained layout checkpoints are saved under outputs/ckpt/.


Inference

Run layout inference on selected scenes with:

python scripts/infer.py \
  --checkpoint /path/to/checkpoint.pth \
  --tokenizer /path/to/tokenizer \
  --dataset_type imaginarium \
  --scene_list '["your/scene/id"]' \
  --output_dir outputs/infer/runs

Results (composed scene .glb and layout metadata) are written under outputs/infer/runs/.


Checklist

  • ✅ Release NaLA training and inference code
  • ⏳ Release NaLA checkpoints (planned for September)
  • ⏳ Release NaLA training data (planned for September)

Citation

If you find this work useful, please cite:

@inproceedings{wan2026nala,
  title={NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation},
  author={Wan, Cheng and Mao, Yongsen and Wu, Wenzheng and Xie, Yuxuan and Xiang, Chucheng and Wang, Runze and Zhang, Xiang and Liu, Zhongyuan and Dai, Rushi and Liu, Yuan},
  booktitle={ECCV},
  year={2026}
}

About

No description, website, or topics provided.

Resources

Stars

33 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages