Oficial implementation of "Point Cloud as a Foreign Language for Multi-modal Large Language Model", accepted at CVPR26.
Python
21
5 commits
updated Mar 25, 2026
Official repository for the paper "Point Cloud as a Foreign Language for Multi-modal Large Language Model"
Multi-modal large language models (MLLMs) have shown remarkable progress in integrating visual and linguistic understanding. Recent efforts have extended these capabilities to 3D understanding through encoder-based architectures that rely on pre-trained 3D encoders to extract geometric features. However, such approaches suffer from semantic misalignment between geometric and linguistic spaces, resolution sensitivity, and substantial computational overhead. In this work, we present SAGE, the first end-to-end 3D MLLM that directly processes raw point clouds without relying on a pre-trained 3D encoder. Our approach introduces a lightweight 3D tokenizer that combines geometric sampling and neighbourhood aggregation with vector quantization to convert point clouds into discrete tokens—treating 3D data as a foreign language that naturally extends the LLM’s vocabulary. Furthermore, to enhance the model’s reasoning capability on complex 3D tasks, we propose a preference optimization training strategy with a semantic alignment–based reward, specifically designed for open-ended 3D question answering where responses are descriptive. Extensive experiments across diverse 3D understanding benchmarks demonstrate that our end-to-end approach outperforms existing encoder-based methods while offering significant advantages in computational efficiency, generalization across LLM backbones, and robustness to input resolution variations.
To start:
cd SAGE
conda create -n SAGE python=3.10 -y
conda activate SAGE
pip install --upgrade pip # enable PEP 660 support
pip install -e .
# * for training
pip install ninja
pip install flash-attn
# * for chamfer_dist
git clone https://github.com/Pang-Yatian/Point-MAE.git
cd ./extensions/chamfer_dist
python setup.py install --user
8192_npy containing 660K point cloud files named {Objaverse_ID}_8192.npy. Each file is a numpy array with dimensions (8192, 6), where the first three dimensions are xyz and the last three dimensions are rgb in [0, 1] range.cat Objaverse_660K_8192_npy_split_a* > Objaverse_660K_8192_npy.tar.gz
tar -xvf Objaverse_660K_8192_npy.tar.gz
SAGE folder, create a folder data and create a soft link to the uncompressed file in the directory.cd SAGE
mkdir data
ln -s /path/to/8192_npy data/objaverse_data
SAGE/data folder, create a directory named anno_data.anno_data directory. The directory should look like this:SAGE/data/anno_data
├── PointLLM_brief_description_660K_filtered.json
├── PointLLM_brief_description_660K.json
└── PointLLM_complex_instruction_70K.json
PointLLM_brief_description_660K_filtered.json is filtered from PointLLM_brief_description_660K.json by removing the 3000 objects we reserved as the validation set.PointLLM_brief_description_val_200_GT.json we use for the benchmarks on Objaverse dataset here, and put it in SAGE/data/anno_data.SAGE folder, create a directory named checkpoints.checkpoints directory.dir_path=.
model_name_or_path=./checkpoints/PointLLM_7B_v1.1_init
data_path=./data/objaverse_data
anno_path=./data/anno_data/PointLLM_brief_description_660K_filtered.json # or PointLLM_brief_description_660K.json (including val sets)
output_dir=./outputs/Train_stage1/$filename
CUDA_VISIBLE_DEVICES=0,1,2,3 torchrun --nnodes=1 --nproc_per_node=4 --master_port=$master_port pointllm/train/train_mem.py \
--model_name_or_path $model_name_or_path \
--data_path $data_path \
--anno_path $anno_path \
--output_dir $output_dir \
--version v1 \
--model_max_length 2048 \
--num_train_epochs 3 \
--per_device_train_batch_size 16 \
--per_device_eval_batch_size 4 \
--gradient_accumulation_steps 2 \
--evaluation_strategy "no" \
--save_strategy "steps" \
--save_steps 100 \
--save_total_limit 1 \
--learning_rate 4e-4 \
--weight_decay 0. \
--warmup_ratio 0.03 \
--lr_scheduler_type "cosine" \
--logging_steps 1 \
--bf16 True \
--fix_llm True \
--fix_pointnet False \
--gradient_checkpointing True \
--report_to tensorboard \
--run_name $filename \
--mm_projector_lr 4e-4 \
--vision_tower_lr 4e-4 \
--group_size 81 \
--num_stages 1 \
--embed_dim 288 \
--LGA_dim 2 2 2 \
--input_points 1024 \
--tune_layer 4 \
--use_color True \
--recon_fp 0 \
--mae_fp 1 \
--mask_dim 4096 \
--mask_ratio 0.0 \
--mae_feature 0 \
--recon_feature 0 \
--pos_embed_mae 0 \
--pos_embed_dim 4096 \
--recon_pos 1 \
--alpha 1.0 \
--beta 0.0 \
--gamma 0.0 \
--vq_cost 0.5 \
--codebook_size 8192 \
--commitment_cost 0.25 \
echo "End of training..."
dir_path=.
model_name_or_path=./outputs/Train_stage1/$filename/checkpoint-15400 # Path to the output dir of stage 1 training
data_path=./data/objaverse_data
anno_path=./data/anno_data/PointLLM_complex_instruction_70K.json
output_dir=./outputs/Train_stage2/$filename
CUDA_VISIBLE_DEVICES=0,1,2,3 torchrun --nnodes=1 --nproc_per_node=4 --master_port=$master_port pointllm/train/train_mem.py \
--model_name_or_path $model_name_or_path \
--data_path $data_path \
--anno_path $anno_path \
--output_dir $output_dir \
--version v1 \
--model_max_length 2048 \
--num_train_epochs 3 \
--per_device_train_batch_size 4 \
--per_device_eval_batch_size 1 \
--gradient_accumulation_steps 2 \
--evaluation_strategy "no" \
--eval_steps 100 \
--save_strategy "steps" \
--save_steps 10000 \
--save_total_limit 1 \
--learning_rate 2e-5 \
--weight_decay 0. \
--warmup_ratio 0.03 \
--lr_scheduler_type "cosine" \
--logging_steps 1 \
--bf16 True \
--fix_llm False \
--fix_pointnet False \
--report_to tensorboard \
--run_name $filename \
--gradient_checkpointing True \
--stage_2 True \
--fsdp "full_shard auto_wrap" \
--fsdp_transformer_layer_cls_to_wrap 'LlamaDecoderLayer' \
--conversation_types "detailed_description" "single_round" "multi_round" \
--mm_projector_lr 2e-5 \
--vision_tower_lr 2e-5 \
--group_size 81 \
--num_stages 1 \
--embed_dim 288 \
--LGA_dim 2 2 2 \
--input_points 1024 \
--tune_layer 0 \
--mask_ratio 0.0 \
--recon_fp 0 \
--mae_fp 1 \
--mask_dim 4096 \
--mae_feature 0 \
--recon_feature 0 \
--recon_pos 1 \
--pos_embed_mae 0 \
--pos_embed_dim 4096 \
--use_color True \
--vq_cost 0.5 \
--codebook_size 8192 \
--commitment_cost 0.25 \
# Model and log paths
MODEL_NAME="./outputs/Train_stage2/$filename"
LOG_SUFFIX="eval"
LOG_DIR="./outputs/new_eval_logs"
LOG_EDIR="./outputs/new_eval_logs"
# Object captioning on Objaverse
CUDA_VISIBLE_DEVICES=0 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type captioning --prompt_index 2 &
# Open Vocabulary Classification on Objaverse
CUDA_VISIBLE_DEVICES=1 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type classification --prompt_index 0
### Traditional Evaluation
CUDA_VISIBLE_DEVICES=0 python pointllm/eval/traditional_evaluator.py --results_path ./outputs/Train_stage2/$filename/evaluation/PointLLM_brief_description_val_200_GT_Objaverse_captioning_prompt2.json
# Object captioning on Objaverse
CUDA_VISIBLE_DEVICES=1 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type captioning --prompt_index 2 > $LOG_EDIR/try_obj_${LOG_SUFFIX}.log 2>&1 &
# Open Vocabulary Classification on Objaverse
CUDA_VISIBLE_DEVICES=2 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type classification --prompt_index 0 > $LOG_EDIR/try_objcls_${LOG_SUFFIX}.log 2>&1 &
{model_name}/evaluation as a dict with the following format:{
"prompt": "",
"results": [
{
"object_id": "",
"ground_truth": "",
"model_output": "",
"label_name": "" # only for classification on modelnet40
}
]
}
This work is under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Python
97.4%
Shell
2.6%
Oficial implementation of "Point Cloud as a Foreign Language for Multi-modal Large Language Model", accepted at CVPR26.
Python
21
5 commits
updated Mar 25, 2026
Official repository for the paper "Point Cloud as a Foreign Language for Multi-modal Large Language Model"
Multi-modal large language models (MLLMs) have shown remarkable progress in integrating visual and linguistic understanding. Recent efforts have extended these capabilities to 3D understanding through encoder-based architectures that rely on pre-trained 3D encoders to extract geometric features. However, such approaches suffer from semantic misalignment between geometric and linguistic spaces, resolution sensitivity, and substantial computational overhead. In this work, we present SAGE, the first end-to-end 3D MLLM that directly processes raw point clouds without relying on a pre-trained 3D encoder. Our approach introduces a lightweight 3D tokenizer that combines geometric sampling and neighbourhood aggregation with vector quantization to convert point clouds into discrete tokens—treating 3D data as a foreign language that naturally extends the LLM’s vocabulary. Furthermore, to enhance the model’s reasoning capability on complex 3D tasks, we propose a preference optimization training strategy with a semantic alignment–based reward, specifically designed for open-ended 3D question answering where responses are descriptive. Extensive experiments across diverse 3D understanding benchmarks demonstrate that our end-to-end approach outperforms existing encoder-based methods while offering significant advantages in computational efficiency, generalization across LLM backbones, and robustness to input resolution variations.
To start:
cd SAGE
conda create -n SAGE python=3.10 -y
conda activate SAGE
pip install --upgrade pip # enable PEP 660 support
pip install -e .
# * for training
pip install ninja
pip install flash-attn
# * for chamfer_dist
git clone https://github.com/Pang-Yatian/Point-MAE.git
cd ./extensions/chamfer_dist
python setup.py install --user
8192_npy containing 660K point cloud files named {Objaverse_ID}_8192.npy. Each file is a numpy array with dimensions (8192, 6), where the first three dimensions are xyz and the last three dimensions are rgb in [0, 1] range.cat Objaverse_660K_8192_npy_split_a* > Objaverse_660K_8192_npy.tar.gz
tar -xvf Objaverse_660K_8192_npy.tar.gz
SAGE folder, create a folder data and create a soft link to the uncompressed file in the directory.cd SAGE
mkdir data
ln -s /path/to/8192_npy data/objaverse_data
SAGE/data folder, create a directory named anno_data.anno_data directory. The directory should look like this:SAGE/data/anno_data
├── PointLLM_brief_description_660K_filtered.json
├── PointLLM_brief_description_660K.json
└── PointLLM_complex_instruction_70K.json
PointLLM_brief_description_660K_filtered.json is filtered from PointLLM_brief_description_660K.json by removing the 3000 objects we reserved as the validation set.PointLLM_brief_description_val_200_GT.json we use for the benchmarks on Objaverse dataset here, and put it in SAGE/data/anno_data.SAGE folder, create a directory named checkpoints.checkpoints directory.dir_path=.
model_name_or_path=./checkpoints/PointLLM_7B_v1.1_init
data_path=./data/objaverse_data
anno_path=./data/anno_data/PointLLM_brief_description_660K_filtered.json # or PointLLM_brief_description_660K.json (including val sets)
output_dir=./outputs/Train_stage1/$filename
CUDA_VISIBLE_DEVICES=0,1,2,3 torchrun --nnodes=1 --nproc_per_node=4 --master_port=$master_port pointllm/train/train_mem.py \
--model_name_or_path $model_name_or_path \
--data_path $data_path \
--anno_path $anno_path \
--output_dir $output_dir \
--version v1 \
--model_max_length 2048 \
--num_train_epochs 3 \
--per_device_train_batch_size 16 \
--per_device_eval_batch_size 4 \
--gradient_accumulation_steps 2 \
--evaluation_strategy "no" \
--save_strategy "steps" \
--save_steps 100 \
--save_total_limit 1 \
--learning_rate 4e-4 \
--weight_decay 0. \
--warmup_ratio 0.03 \
--lr_scheduler_type "cosine" \
--logging_steps 1 \
--bf16 True \
--fix_llm True \
--fix_pointnet False \
--gradient_checkpointing True \
--report_to tensorboard \
--run_name $filename \
--mm_projector_lr 4e-4 \
--vision_tower_lr 4e-4 \
--group_size 81 \
--num_stages 1 \
--embed_dim 288 \
--LGA_dim 2 2 2 \
--input_points 1024 \
--tune_layer 4 \
--use_color True \
--recon_fp 0 \
--mae_fp 1 \
--mask_dim 4096 \
--mask_ratio 0.0 \
--mae_feature 0 \
--recon_feature 0 \
--pos_embed_mae 0 \
--pos_embed_dim 4096 \
--recon_pos 1 \
--alpha 1.0 \
--beta 0.0 \
--gamma 0.0 \
--vq_cost 0.5 \
--codebook_size 8192 \
--commitment_cost 0.25 \
echo "End of training..."
dir_path=.
model_name_or_path=./outputs/Train_stage1/$filename/checkpoint-15400 # Path to the output dir of stage 1 training
data_path=./data/objaverse_data
anno_path=./data/anno_data/PointLLM_complex_instruction_70K.json
output_dir=./outputs/Train_stage2/$filename
CUDA_VISIBLE_DEVICES=0,1,2,3 torchrun --nnodes=1 --nproc_per_node=4 --master_port=$master_port pointllm/train/train_mem.py \
--model_name_or_path $model_name_or_path \
--data_path $data_path \
--anno_path $anno_path \
--output_dir $output_dir \
--version v1 \
--model_max_length 2048 \
--num_train_epochs 3 \
--per_device_train_batch_size 4 \
--per_device_eval_batch_size 1 \
--gradient_accumulation_steps 2 \
--evaluation_strategy "no" \
--eval_steps 100 \
--save_strategy "steps" \
--save_steps 10000 \
--save_total_limit 1 \
--learning_rate 2e-5 \
--weight_decay 0. \
--warmup_ratio 0.03 \
--lr_scheduler_type "cosine" \
--logging_steps 1 \
--bf16 True \
--fix_llm False \
--fix_pointnet False \
--report_to tensorboard \
--run_name $filename \
--gradient_checkpointing True \
--stage_2 True \
--fsdp "full_shard auto_wrap" \
--fsdp_transformer_layer_cls_to_wrap 'LlamaDecoderLayer' \
--conversation_types "detailed_description" "single_round" "multi_round" \
--mm_projector_lr 2e-5 \
--vision_tower_lr 2e-5 \
--group_size 81 \
--num_stages 1 \
--embed_dim 288 \
--LGA_dim 2 2 2 \
--input_points 1024 \
--tune_layer 0 \
--mask_ratio 0.0 \
--recon_fp 0 \
--mae_fp 1 \
--mask_dim 4096 \
--mae_feature 0 \
--recon_feature 0 \
--recon_pos 1 \
--pos_embed_mae 0 \
--pos_embed_dim 4096 \
--use_color True \
--vq_cost 0.5 \
--codebook_size 8192 \
--commitment_cost 0.25 \
# Model and log paths
MODEL_NAME="./outputs/Train_stage2/$filename"
LOG_SUFFIX="eval"
LOG_DIR="./outputs/new_eval_logs"
LOG_EDIR="./outputs/new_eval_logs"
# Object captioning on Objaverse
CUDA_VISIBLE_DEVICES=0 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type captioning --prompt_index 2 &
# Open Vocabulary Classification on Objaverse
CUDA_VISIBLE_DEVICES=1 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type classification --prompt_index 0
### Traditional Evaluation
CUDA_VISIBLE_DEVICES=0 python pointllm/eval/traditional_evaluator.py --results_path ./outputs/Train_stage2/$filename/evaluation/PointLLM_brief_description_val_200_GT_Objaverse_captioning_prompt2.json
# Object captioning on Objaverse
CUDA_VISIBLE_DEVICES=1 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type captioning --prompt_index 2 > $LOG_EDIR/try_obj_${LOG_SUFFIX}.log 2>&1 &
# Open Vocabulary Classification on Objaverse
CUDA_VISIBLE_DEVICES=2 python pointllm/eval/eval_objaverse.py --model_name $MODEL_NAME --task_type classification --prompt_index 0 > $LOG_EDIR/try_objcls_${LOG_SUFFIX}.log 2>&1 &
{model_name}/evaluation as a dict with the following format:{
"prompt": "",
"results": [
{
"object_id": "",
"ground_truth": "",
"model_output": "",
"label_name": "" # only for classification on modelnet40
}
]
}
This work is under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Python
97.4%
Shell
2.6%