PKU-ICST-MIPL/TARA_CVPR2026

Python

20

6 commits

updated Mar 21, 2026

See the code

README

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

Hulingxiao He · Zhi Tan · Yuxin Peng

CVPR 2026

Paper | Data

🔥 News

  • Mar 2026: 🌼🌼🌼 Code is available now. Welcome to follow our work and give us a star 🌟!
  • Feb 2026: 🎉🎉🎉 TARA is accepted to CVPR 2026! See you in Denver this June!
  • Fine-R1 (ICLR 2026): the first MLLM to surpass various strong CLIP-like models (e.g., SigLIP-L) in FGVR. 【Paper】【Code】
  • Finedefics (ICLR 2025): revisiting three quintessential capabilities of MLLMs for FGVR and position of the root cause as a misalignment problem. 【Paper】【Code】

🌟 Motivation

Large Multimodal Models struggle with hierarchical visual recognition (HVR), failing to obey the hierarchical consistency on both known and novel categories.

Overview

📖 Methodology

We propose Taxonomy-Aware Representation Alignment (TARA), a simple but effective framework that explicitly aligns intermediate representations of LMMs with visual and text features from pretrained Biology Foundation Models (BFMs), thereby injecting taxonomic knowledge and enabling richer, hierarchy-aware visual recognition.

Method

Main Results

Experiments demonstrate that TARA consistently enhances LMMs’ hierarchical consistency and leaf node accuracy, enabling reliable recognition of both known and novel categories within complex biological taxonomies.

Main Results

🛠️ Installation

cd CLS-RL
conda create -n cls-rl python=3.11
conda activate cls-rl
bash setup.sh
pip install flash-attn==2.7.2.post1 --no-build-isolation
(Optional) export HF_ENDPOINT=https://hf-mirror.com

🔥 Training

Data Preparation

Download 1-shot training data of iNaturalist-2021 from Hugging Face and put it under the directory CLS-RL/data.

No-Thinking RFT (baseline)
# Qwen3-VL-2B
bash fewshot_no-think-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-qwen2_5.sh
No-Thinking RFT + TARA (ours)
# Qwen3-VL-2B
bash fewshot_no-think-tara-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-tara-qwen2_5.sh

📋 Evaluation

Data Preparation

Update image paths in the JSON files LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_animalia_with_similar_choice_new_sample1.jsonl and LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_plantae_with_similar_choice_new_sample1.jsonl to point to your image directory. Note that we randomly sample 1-shot data for fast evaluation in our work.

No-Thinking RFT (baseline)
cd LLM-Hierarchical-Consistency
bash scripts/baselines.sh
No-Thinking RFT + TARA (ours)
cd LLM-Hierarchical-Consistency
bash scripts/ours.sh

🥰 Acknowledgements

We thank the CLS-RL, LLM-Hierarchical-Consistency, and BioCLIP2 for providing the foundational codebase that we adapted to implement TARA.

📝 Citation

If you find it useful for your research and applications, please cite related papers using this BibTeX:

@article{he2026taxonomy,
  title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
  author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
  journal={arXiv preprint arXiv:2603.00431},
  year={2026}
}

@article{he2026fine,
  title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
  author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
  journal={arXiv preprint arXiv:2602.07605},
  year={2026}
}

@article{he2025analyzing,
  title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
  author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
  journal={arXiv preprint arXiv:2501.15140},
  year={2025}
}

📄 License

This project is licensed under the MIT License.


PKU-ICST-MIPL/TARA_CVPR2026

Python

20

6 commits

updated Mar 21, 2026

See the code

README

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

Hulingxiao He · Zhi Tan · Yuxin Peng

CVPR 2026

Paper | Data

🔥 News

  • Mar 2026: 🌼🌼🌼 Code is available now. Welcome to follow our work and give us a star 🌟!
  • Feb 2026: 🎉🎉🎉 TARA is accepted to CVPR 2026! See you in Denver this June!
  • Fine-R1 (ICLR 2026): the first MLLM to surpass various strong CLIP-like models (e.g., SigLIP-L) in FGVR. 【Paper】【Code】
  • Finedefics (ICLR 2025): revisiting three quintessential capabilities of MLLMs for FGVR and position of the root cause as a misalignment problem. 【Paper】【Code】

🌟 Motivation

Large Multimodal Models struggle with hierarchical visual recognition (HVR), failing to obey the hierarchical consistency on both known and novel categories.

Overview

📖 Methodology

We propose Taxonomy-Aware Representation Alignment (TARA), a simple but effective framework that explicitly aligns intermediate representations of LMMs with visual and text features from pretrained Biology Foundation Models (BFMs), thereby injecting taxonomic knowledge and enabling richer, hierarchy-aware visual recognition.

Method

Main Results

Experiments demonstrate that TARA consistently enhances LMMs’ hierarchical consistency and leaf node accuracy, enabling reliable recognition of both known and novel categories within complex biological taxonomies.

Main Results

🛠️ Installation

cd CLS-RL
conda create -n cls-rl python=3.11
conda activate cls-rl
bash setup.sh
pip install flash-attn==2.7.2.post1 --no-build-isolation
(Optional) export HF_ENDPOINT=https://hf-mirror.com

🔥 Training

Data Preparation

Download 1-shot training data of iNaturalist-2021 from Hugging Face and put it under the directory CLS-RL/data.

No-Thinking RFT (baseline)
# Qwen3-VL-2B
bash fewshot_no-think-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-qwen2_5.sh
No-Thinking RFT + TARA (ours)
# Qwen3-VL-2B
bash fewshot_no-think-tara-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-tara-qwen2_5.sh

📋 Evaluation

Data Preparation

Update image paths in the JSON files LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_animalia_with_similar_choice_new_sample1.jsonl and LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_plantae_with_similar_choice_new_sample1.jsonl to point to your image directory. Note that we randomly sample 1-shot data for fast evaluation in our work.

No-Thinking RFT (baseline)
cd LLM-Hierarchical-Consistency
bash scripts/baselines.sh
No-Thinking RFT + TARA (ours)
cd LLM-Hierarchical-Consistency
bash scripts/ours.sh

🥰 Acknowledgements

We thank the CLS-RL, LLM-Hierarchical-Consistency, and BioCLIP2 for providing the foundational codebase that we adapted to implement TARA.

📝 Citation

If you find it useful for your research and applications, please cite related papers using this BibTeX:

@article{he2026taxonomy,
  title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
  author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
  journal={arXiv preprint arXiv:2603.00431},
  year={2026}
}

@article{he2026fine,
  title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
  author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
  journal={arXiv preprint arXiv:2602.07605},
  year={2026}
}

@article{he2025analyzing,
  title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
  author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
  journal={arXiv preprint arXiv:2501.15140},
  year={2025}
}

📄 License

This project is licensed under the MIT License.