Hulingxiao He · Zijun Geng · Yuxin Peng
Fine-R1 generates Chain-of-Thought (CoT) before producing the final fine-grained visual recognition (FGVR) answer. It utilizes CoT supervised fine-tuning (SFT) and Triplet Augmented Policy Optimization (TAPO), learning the reasoning process with only few-shot samples per category.
(1) Closed-world evaluation: In comparison to general and reasoning MLLMs, and contrastive CLIP models, Fine-R1 excels in identifying both seen and unseen categories:
(2) Open-world evaluation: Fine-R1 establishes new state-of-the-art performance with only 4-shot training samples per sub-category, achieving 74.80% relative semantic similarity on average:
Please follow the instructions to set up the training environment: https://github.com/hiyouga/LLaMA-Factory. We recommend creating a new conda environment specifically for this stage to avoid potential package version conflicts.
FineR1_ICLR2026/LLaMA-Factory/data/Fine-R1-Stage1-data.json to point to your image directory.# 3B model
cd LLaMA-Factory
bash cot_sft_3b.sh
# 7B model
cd LLaMA-Factory
bash cot_sft_7b.sh
conda create -n tapo python=3.10
conda activate tapo
cd TAPO
bash scripts/install.sh
conda create -n tapo python=3.10
conda activate tapo
cd TAPO
pip install -e .
pip install flash-attn==2.7.3 --no-build-isolation
Download the training data of stage 2 from Fine-R1-Stage2-data and unzip it in FineR1_ICLR2026/TAPO/data directory. You can also refer to the data format of Fine-R1-Stage2-data to create your own customized dataset. Note that the image_noisy and image_aug should be preprocessed to the same size with the original images.
The main training pipeline is adopted from EasyR1. We support training with different configurations for both Fine-R1-3B and Fine-R1-7B models:
# 3B model
cd TAPO
bash examples/tapo/tapo_3b.sh
# 7B model
cd TAPO
bash examples/tapo/tapo_7b.sh
We evaluate the models in both closed-world (multi-choice) and open-world (question-answering) FGVR, on both seen and unseen categories. Before evaluation, update image paths image_path in the JSONL files FineR1_ICLR2026/eval/data/*.jsonl to point to your image directory.
# Closed-world evaluation
cd eval
bash scripts/eval_closed.sh
# Open-world evaluation
cd eval
bash scripts/eval_open.sh
We thank the PAPO, NoisyRollout, LlamaFactory, and EasyR1 team for providing the foundational codebase that we adapted to implement Fine-R1.
If you find it useful for your research and applications, please cite related papers using this BibTeX:
@article{he2026fine,
title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
journal={arXiv preprint arXiv:2602.07605},
year={2026}
}
@article{he2026taxonomy,
title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
journal={arXiv preprint arXiv:2603.00431},
year={2026}
}
@article{he2025analyzing,
title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
journal={arXiv preprint arXiv:2501.15140},
year={2025}
}
This project is licensed under the MIT License.
Hulingxiao He · Zijun Geng · Yuxin Peng
Fine-R1 generates Chain-of-Thought (CoT) before producing the final fine-grained visual recognition (FGVR) answer. It utilizes CoT supervised fine-tuning (SFT) and Triplet Augmented Policy Optimization (TAPO), learning the reasoning process with only few-shot samples per category.
(1) Closed-world evaluation: In comparison to general and reasoning MLLMs, and contrastive CLIP models, Fine-R1 excels in identifying both seen and unseen categories:
(2) Open-world evaluation: Fine-R1 establishes new state-of-the-art performance with only 4-shot training samples per sub-category, achieving 74.80% relative semantic similarity on average:
Please follow the instructions to set up the training environment: https://github.com/hiyouga/LLaMA-Factory. We recommend creating a new conda environment specifically for this stage to avoid potential package version conflicts.
FineR1_ICLR2026/LLaMA-Factory/data/Fine-R1-Stage1-data.json to point to your image directory.# 3B model
cd LLaMA-Factory
bash cot_sft_3b.sh
# 7B model
cd LLaMA-Factory
bash cot_sft_7b.sh
conda create -n tapo python=3.10
conda activate tapo
cd TAPO
bash scripts/install.sh
conda create -n tapo python=3.10
conda activate tapo
cd TAPO
pip install -e .
pip install flash-attn==2.7.3 --no-build-isolation
Download the training data of stage 2 from Fine-R1-Stage2-data and unzip it in FineR1_ICLR2026/TAPO/data directory. You can also refer to the data format of Fine-R1-Stage2-data to create your own customized dataset. Note that the image_noisy and image_aug should be preprocessed to the same size with the original images.
The main training pipeline is adopted from EasyR1. We support training with different configurations for both Fine-R1-3B and Fine-R1-7B models:
# 3B model
cd TAPO
bash examples/tapo/tapo_3b.sh
# 7B model
cd TAPO
bash examples/tapo/tapo_7b.sh
We evaluate the models in both closed-world (multi-choice) and open-world (question-answering) FGVR, on both seen and unseen categories. Before evaluation, update image paths image_path in the JSONL files FineR1_ICLR2026/eval/data/*.jsonl to point to your image directory.
# Closed-world evaluation
cd eval
bash scripts/eval_closed.sh
# Open-world evaluation
cd eval
bash scripts/eval_open.sh
We thank the PAPO, NoisyRollout, LlamaFactory, and EasyR1 team for providing the foundational codebase that we adapted to implement Fine-R1.
If you find it useful for your research and applications, please cite related papers using this BibTeX:
@article{he2026fine,
title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
journal={arXiv preprint arXiv:2602.07605},
year={2026}
}
@article{he2026taxonomy,
title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
journal={arXiv preprint arXiv:2603.00431},
year={2026}
}
@article{he2025analyzing,
title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
journal={arXiv preprint arXiv:2501.15140},
year={2025}
}
This project is licensed under the MIT License.