PKU-ICST-MIPL/FineR1_ICLR2026

Python

78

10 commits

updated Apr 4, 2026

See the code

README

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

Hulingxiao He · Zijun Geng · Yuxin Peng

ICLR 2026

Models (Downloads: 23.7k🔥) | Paper | Data | Poster | Slides | Video

🔥 New

  • Apr 2026: 🌼🌼🌼 Poster, slides, and video are available now.
  • Feb 2026: 🌼🌼🌼 Checkpoints after Stage 1 are newly released here. Download and use them for your customized Stage 2 traini
  • Feb 2026: 🌼🌼🌼 Code, data, and models are released now. Welcome to follow our work and give us a star 🌟!
  • Jan 2026: 🎉🎉🎉 Fine-R1 is accepted to ICLR 2026! See you in Rio de Janeiro this April!
  • TARA (CVPR 2026): using fine-grained category tree to boost hierarchical visual recognition capability of MLLMs. 【Paper】【Code】
  • Finedefics (ICLR 2025): revisiting three quintessential capabilities of MLLMs for FGVR and position of the root cause as a misalignment problem. 【Paper】【Code】

🌟 Key Highlights

  • To the best of our knowledge, our Fine-R1 is the first MLLM to surpass various strong CLIP-like models (e.g., SigLIP-L) in FGVR: It's widely acknowledged that general MLLMs underperform contrastive models like CLIP/SigLIP on fine-grained tasks. Our work bridges this gap and strongly indicates the potential of generative MLLMs for discriminative vision tasks.
Overview

📖 Methodology

Fine-R1 generates Chain-of-Thought (CoT) before producing the final fine-grained visual recognition (FGVR) answer. It utilizes CoT supervised fine-tuning (SFT) and Triplet Augmented Policy Optimization (TAPO), learning the reasoning process with only few-shot samples per category.

Method

Main Results

(1) Closed-world evaluation: In comparison to general and reasoning MLLMs, and contrastive CLIP models, Fine-R1 excels in identifying both seen and unseen categories:

Main Results

(2) Open-world evaluation: Fine-R1 establishes new state-of-the-art performance with only 4-shot training samples per sub-category, achieving 74.80% relative semantic similarity on average:

Main Results

🚀 Quick Start

Stage1: CoT SFT

(1) Environment Setup

Please follow the instructions to set up the training environment: https://github.com/hiyouga/LLaMA-Factory. We recommend creating a new conda environment specifically for this stage to avoid potential package version conflicts.

(2) Data Preparation

  1. Download and prepare the training images of 6 FGVR datasets:
  • FGVC-Aircraft ✈️
  • CaltechUCSD Bird-200 🦤
  • Stanford Car-196 🚙
  • Stanford Dog-120 🐕
  • Flower-102 🌼
  • Oxford-IIIT Pet-37 🐈
  1. Update image paths in the JSON files FineR1_ICLR2026/LLaMA-Factory/data/Fine-R1-Stage1-data.json to point to your image directory.

(3) Training

# 3B model
cd LLaMA-Factory
bash cot_sft_3b.sh

# 7B model  
cd LLaMA-Factory
bash cot_sft_7b.sh

Stage2: TAPO

(1) Environment Setup

Option 1: All-in-one Installation Script

conda create -n tapo python=3.10
conda activate tapo

cd TAPO
bash scripts/install.sh

Option 2: Using pip

conda create -n tapo python=3.10
conda activate tapo

cd TAPO
pip install -e .
pip install flash-attn==2.7.3 --no-build-isolation

(2) Data Preparation

Download the training data of stage 2 from Fine-R1-Stage2-data and unzip it in FineR1_ICLR2026/TAPO/data directory. You can also refer to the data format of Fine-R1-Stage2-data to create your own customized dataset. Note that the image_noisy and image_aug should be preprocessed to the same size with the original images.

(3) Training

The main training pipeline is adopted from EasyR1. We support training with different configurations for both Fine-R1-3B and Fine-R1-7B models:

# 3B model
cd TAPO
bash examples/tapo/tapo_3b.sh

# 7B model  
cd TAPO
bash examples/tapo/tapo_7b.sh

Performance Evaluation

We evaluate the models in both closed-world (multi-choice) and open-world (question-answering) FGVR, on both seen and unseen categories. Before evaluation, update image paths image_path in the JSONL files FineR1_ICLR2026/eval/data/*.jsonl to point to your image directory.

# Closed-world evaluation
cd eval
bash scripts/eval_closed.sh

# Open-world evaluation
cd eval
bash scripts/eval_open.sh

🥰 Acknowledgements

We thank the PAPO, NoisyRollout, LlamaFactory, and EasyR1 team for providing the foundational codebase that we adapted to implement Fine-R1.

📝 Citation

If you find it useful for your research and applications, please cite related papers using this BibTeX:

@article{he2026fine,
  title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
  author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
  journal={arXiv preprint arXiv:2602.07605},
  year={2026}
}

@article{he2026taxonomy,
  title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
  author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
  journal={arXiv preprint arXiv:2603.00431},
  year={2026}
}

@article{he2025analyzing,
  title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
  author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
  journal={arXiv preprint arXiv:2501.15140},
  year={2025}
}

📄 License

This project is licensed under the MIT License.


PKU-ICST-MIPL/FineR1_ICLR2026

Python

78

10 commits

updated Apr 4, 2026

See the code

README

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

Hulingxiao He · Zijun Geng · Yuxin Peng

ICLR 2026

Models (Downloads: 23.7k🔥) | Paper | Data | Poster | Slides | Video

🔥 New

  • Apr 2026: 🌼🌼🌼 Poster, slides, and video are available now.
  • Feb 2026: 🌼🌼🌼 Checkpoints after Stage 1 are newly released here. Download and use them for your customized Stage 2 traini
  • Feb 2026: 🌼🌼🌼 Code, data, and models are released now. Welcome to follow our work and give us a star 🌟!
  • Jan 2026: 🎉🎉🎉 Fine-R1 is accepted to ICLR 2026! See you in Rio de Janeiro this April!
  • TARA (CVPR 2026): using fine-grained category tree to boost hierarchical visual recognition capability of MLLMs. 【Paper】【Code】
  • Finedefics (ICLR 2025): revisiting three quintessential capabilities of MLLMs for FGVR and position of the root cause as a misalignment problem. 【Paper】【Code】

🌟 Key Highlights

  • To the best of our knowledge, our Fine-R1 is the first MLLM to surpass various strong CLIP-like models (e.g., SigLIP-L) in FGVR: It's widely acknowledged that general MLLMs underperform contrastive models like CLIP/SigLIP on fine-grained tasks. Our work bridges this gap and strongly indicates the potential of generative MLLMs for discriminative vision tasks.
Overview

📖 Methodology

Fine-R1 generates Chain-of-Thought (CoT) before producing the final fine-grained visual recognition (FGVR) answer. It utilizes CoT supervised fine-tuning (SFT) and Triplet Augmented Policy Optimization (TAPO), learning the reasoning process with only few-shot samples per category.

Method

Main Results

(1) Closed-world evaluation: In comparison to general and reasoning MLLMs, and contrastive CLIP models, Fine-R1 excels in identifying both seen and unseen categories:

Main Results

(2) Open-world evaluation: Fine-R1 establishes new state-of-the-art performance with only 4-shot training samples per sub-category, achieving 74.80% relative semantic similarity on average:

Main Results

🚀 Quick Start

Stage1: CoT SFT

(1) Environment Setup

Please follow the instructions to set up the training environment: https://github.com/hiyouga/LLaMA-Factory. We recommend creating a new conda environment specifically for this stage to avoid potential package version conflicts.

(2) Data Preparation

  1. Download and prepare the training images of 6 FGVR datasets:
  • FGVC-Aircraft ✈️
  • CaltechUCSD Bird-200 🦤
  • Stanford Car-196 🚙
  • Stanford Dog-120 🐕
  • Flower-102 🌼
  • Oxford-IIIT Pet-37 🐈
  1. Update image paths in the JSON files FineR1_ICLR2026/LLaMA-Factory/data/Fine-R1-Stage1-data.json to point to your image directory.

(3) Training

# 3B model
cd LLaMA-Factory
bash cot_sft_3b.sh

# 7B model  
cd LLaMA-Factory
bash cot_sft_7b.sh

Stage2: TAPO

(1) Environment Setup

Option 1: All-in-one Installation Script

conda create -n tapo python=3.10
conda activate tapo

cd TAPO
bash scripts/install.sh

Option 2: Using pip

conda create -n tapo python=3.10
conda activate tapo

cd TAPO
pip install -e .
pip install flash-attn==2.7.3 --no-build-isolation

(2) Data Preparation

Download the training data of stage 2 from Fine-R1-Stage2-data and unzip it in FineR1_ICLR2026/TAPO/data directory. You can also refer to the data format of Fine-R1-Stage2-data to create your own customized dataset. Note that the image_noisy and image_aug should be preprocessed to the same size with the original images.

(3) Training

The main training pipeline is adopted from EasyR1. We support training with different configurations for both Fine-R1-3B and Fine-R1-7B models:

# 3B model
cd TAPO
bash examples/tapo/tapo_3b.sh

# 7B model  
cd TAPO
bash examples/tapo/tapo_7b.sh

Performance Evaluation

We evaluate the models in both closed-world (multi-choice) and open-world (question-answering) FGVR, on both seen and unseen categories. Before evaluation, update image paths image_path in the JSONL files FineR1_ICLR2026/eval/data/*.jsonl to point to your image directory.

# Closed-world evaluation
cd eval
bash scripts/eval_closed.sh

# Open-world evaluation
cd eval
bash scripts/eval_open.sh

🥰 Acknowledgements

We thank the PAPO, NoisyRollout, LlamaFactory, and EasyR1 team for providing the foundational codebase that we adapted to implement Fine-R1.

📝 Citation

If you find it useful for your research and applications, please cite related papers using this BibTeX:

@article{he2026fine,
  title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
  author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
  journal={arXiv preprint arXiv:2602.07605},
  year={2026}
}

@article{he2026taxonomy,
  title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
  author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
  journal={arXiv preprint arXiv:2603.00431},
  year={2026}
}

@article{he2025analyzing,
  title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
  author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
  journal={arXiv preprint arXiv:2501.15140},
  year={2025}
}

📄 License

This project is licensed under the MIT License.