Xiao Luo1
Mingyang Du1
Xin Zhou1
Tianrui Feng1
Xiwu Chen2
Xiaofan Li3
Jiangning Zhang3
Dingkang Liang1,*
1Huazhong University of Science and Technology, China
2Megvii, China
3Zhejiang University, China
{dkliang, tianruifeng, xzhou03}@hust.edu.cn
*Corresponding author
ROAD transfers discriminative 3D semantics into shape generation through global feature alignment and token-level matching. This release contains the code and configurations for:
configs/baseline.yaml: Step1X-3D rectified-flow baseline.configs/uni3d_repa.yaml: ROAD with a frozen Uni3D teacher, global REPA alignment, and token-level Hungarian matching.
Model weights and full training datasets are not included.
The following textured GLBs are rendered as 360-degree turntables.
![]() 001 · Bread man |
![]() 002 · Cat |
![]() 003 · Food bowl |
![]() 004 · Fire hydrant |
![]() 005 · Aircraft |
![]() 006 · Castle |
Run all commands from the repository root.
The training setup uses Python 3.10, PyTorch 2.5.1, and CUDA 11.8.
conda env create -f environment.yml
conda activate step1x
pip install -r requirements.txt
pip install flash-attn==2.8.3 --no-build-isolationUse the existing evaluation environment:
conda activate driveuni3dAlternatively, add the evaluation packages to step1x:
conda activate step1x
pip install -r requirements-eval.txtTraining uses its own Uni3D subset under training/uni3d/; Uni3D-I uses the
separate inference subset under evaluation/uni3d/. Both select point groups
with farthest-point sampling.
Prepare the following files:
pretrained/
├── vae/
│ └── diffusion_pytorch_model.safetensors
├── visual_encoder/
│ ├── config.json
│ ├── model.safetensors
│ └── preprocessor_config.json
├── uni3d/
│ └── model.pt
└── evaluation/
├── uni3d_openclip.bin
├── ulip_openclip.bin
└── ulip2_pointbert.pt
The baseline uses the Step1X-3D VAE and DINOv2 visual encoder. ROAD training
and Uni3D-I additionally use pretrained/uni3d/model.pt. The remaining
evaluation checkpoints are used by Uni3D-I and ULIP-I. See
pretrained/README.md for details.
data/3d_data/
├── train.json
├── val.json
├── test.json
├── surfaces/
│ └── <uid>.npz
└── 3d_images/
└── <uid>/
├── 012.png
├── 013.png
└── ...
Each split is a JSON list of UIDs. Each NPZ file contains surface and
sharp_surface, both stored as N x 6 XYZ-and-normal arrays. Views 12–23 are
used during training. A small synthetic dataset for checking the complete
pipeline is included at data/demo_3d_data.
For details on mesh preprocessing and dataset preparation, please refer to the data preprocessing pipeline provided by Step1X-3D.
Use another dataset root with an OmegaConf override:
GPU_IDS=0 bash scripts/train_baseline.sh data.root_dir=/path/to/3d_dataevaluation_data/
├── manifest.json
├── images/
│ └── <uid>/eval.png
└── meshes/
└── <uid>/eval.glb
manifest.json is a JSON list of UIDs shared by the image and mesh roots.
Nested UIDs such as category/example are supported.
# Single GPU
GPU_IDS=0 bash scripts/train_baseline.sh
# Multiple GPUs on one node
GPU_IDS=0,1,2,3 bash scripts/train_baseline.sh# Single GPU
GPU_IDS=0 bash scripts/train_uni3d_repa.sh
# Multiple GPUs on one node
GPU_IDS=0,1,2,3 bash scripts/train_uni3d_repa.shThe default ROAD configuration uses 10,000 alignment points, global alignment
weight 0.5, token-matching weight 0.1, and starts token matching at epoch 3.
The frozen teacher is configured by configs/uni3d_g.json. To use the SciPy
matcher, append system.matcher=cpu.
All OmegaConf options can be appended to the launcher:
GPU_IDS=0,1 bash scripts/train_uni3d_repa.sh \
data.batch_size=4 \
data.num_workers=4 \
trainer.max_epochs=10Resume from a Lightning or DeepSpeed checkpoint with:
GPU_IDS=0,1 bash scripts/train_uni3d_repa.sh resume=/path/to/last.ckptRun Uni3D-I and ULIP-I on two GPUs:
conda activate driveuni3d
GPU_IDS=0,1 bash scripts/evaluate.sh \
evaluation_data/manifest.json \
evaluation_data/images \
evaluation_data/meshes \
outputs/evaluationThe command writes:
outputs/evaluation/
├── uni3d_i.json
└── ulip_i.json
Run only Uni3D-I on one GPU:
CUDA_VISIBLE_DEVICES=0 python -m evaluation.evaluate_uni3d \
--manifest evaluation_data/manifest.json \
--image-root evaluation_data/images \
--glb-root evaluation_data/meshes \
--openclip-checkpoint pretrained/evaluation/uni3d_openclip.bin \
--output outputs/evaluation/uni3d_i.jsonRun only ULIP-I on one GPU:
CUDA_VISIBLE_DEVICES=0 python -m evaluation.evaluate_ulip \
--manifest evaluation_data/manifest.json \
--image-root evaluation_data/images \
--glb-root evaluation_data/meshes \
--openclip-checkpoint pretrained/evaluation/ulip_openclip.bin \
--ulip-checkpoint pretrained/evaluation/ulip2_pointbert.pt \
--output outputs/evaluation/ulip_i.jsonTraining checkpoints and logs are written below exp_root_dir. Inspect
TensorBoard logs with:
tensorboard --logdir outputsThe Step1X-3D-derived training code is distributed under Apache License 2.0.
The reduced Uni3D subsets under training/uni3d/ and evaluation/uni3d/
retain the upstream MIT license. The reduced ULIP evaluation code retains the
upstream BSD 3-Clause license.
See LICENSE, NOTICE, MODIFICATIONS.md, training/uni3d/LICENSE,
evaluation/uni3d/LICENSE, and evaluation/LICENSE-ULIP. Model weights and
external datasets are distributed separately under their respective terms.
This project builds upon Step1X-3D, Uni3D, and ULIP. We thank the authors of these projects for their contributions to the open-source community.
If you find this repository useful for your research, please consider citing our paper:
@misc{luo2026roadreciprocalobjectivealignmentdiscriminative,
title = {ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation},
author = {Xiao Luo and Mingyang Du and Xin Zhou and Tianrui Feng and Xiwu Chen and Xiaofan Li and Jiangning Zhang and Dingkang Liang},
year = {2026},
eprint = {2607.28581},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2607.28581}
}





