Hulingxiao He · Zhi Tan · Yuxin Peng
Large Multimodal Models struggle with hierarchical visual recognition (HVR), failing to obey the hierarchical consistency on both known and novel categories.
We propose Taxonomy-Aware Representation Alignment (TARA), a simple but effective framework that explicitly aligns intermediate representations of LMMs with visual and text features from pretrained Biology Foundation Models (BFMs), thereby injecting taxonomic knowledge and enabling richer, hierarchy-aware visual recognition.
Experiments demonstrate that TARA consistently enhances LMMs’ hierarchical consistency and leaf node accuracy, enabling reliable recognition of both known and novel categories within complex biological taxonomies.
cd CLS-RL
conda create -n cls-rl python=3.11
conda activate cls-rl
bash setup.sh
pip install flash-attn==2.7.2.post1 --no-build-isolation
(Optional) export HF_ENDPOINT=https://hf-mirror.com
Download 1-shot training data of iNaturalist-2021 from Hugging Face and put it under the directory CLS-RL/data.
# Qwen3-VL-2B
bash fewshot_no-think-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-qwen2_5.sh
# Qwen3-VL-2B
bash fewshot_no-think-tara-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-tara-qwen2_5.sh
Update image paths in the JSON files LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_animalia_with_similar_choice_new_sample1.jsonl and LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_plantae_with_similar_choice_new_sample1.jsonl to point to your image directory. Note that we randomly sample 1-shot data for fast evaluation in our work.
cd LLM-Hierarchical-Consistency
bash scripts/baselines.sh
cd LLM-Hierarchical-Consistency
bash scripts/ours.sh
We thank the CLS-RL, LLM-Hierarchical-Consistency, and BioCLIP2 for providing the foundational codebase that we adapted to implement TARA.
If you find it useful for your research and applications, please cite related papers using this BibTeX:
@article{he2026taxonomy,
title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
journal={arXiv preprint arXiv:2603.00431},
year={2026}
}
@article{he2026fine,
title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
journal={arXiv preprint arXiv:2602.07605},
year={2026}
}
@article{he2025analyzing,
title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
journal={arXiv preprint arXiv:2501.15140},
year={2025}
}
This project is licensed under the MIT License.
Hulingxiao He · Zhi Tan · Yuxin Peng
Large Multimodal Models struggle with hierarchical visual recognition (HVR), failing to obey the hierarchical consistency on both known and novel categories.
We propose Taxonomy-Aware Representation Alignment (TARA), a simple but effective framework that explicitly aligns intermediate representations of LMMs with visual and text features from pretrained Biology Foundation Models (BFMs), thereby injecting taxonomic knowledge and enabling richer, hierarchy-aware visual recognition.
Experiments demonstrate that TARA consistently enhances LMMs’ hierarchical consistency and leaf node accuracy, enabling reliable recognition of both known and novel categories within complex biological taxonomies.
cd CLS-RL
conda create -n cls-rl python=3.11
conda activate cls-rl
bash setup.sh
pip install flash-attn==2.7.2.post1 --no-build-isolation
(Optional) export HF_ENDPOINT=https://hf-mirror.com
Download 1-shot training data of iNaturalist-2021 from Hugging Face and put it under the directory CLS-RL/data.
# Qwen3-VL-2B
bash fewshot_no-think-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-qwen2_5.sh
# Qwen3-VL-2B
bash fewshot_no-think-tara-qwen3.sh
# Qwen2.5-VL-3B
bash fewshot_no-think-tara-qwen2_5.sh
Update image paths in the JSON files LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_animalia_with_similar_choice_new_sample1.jsonl and LLM-Hierarchical-Consistency/data/annotations/similar_choices/inat21_plantae_with_similar_choice_new_sample1.jsonl to point to your image directory. Note that we randomly sample 1-shot data for fast evaluation in our work.
cd LLM-Hierarchical-Consistency
bash scripts/baselines.sh
cd LLM-Hierarchical-Consistency
bash scripts/ours.sh
We thank the CLS-RL, LLM-Hierarchical-Consistency, and BioCLIP2 for providing the foundational codebase that we adapted to implement TARA.
If you find it useful for your research and applications, please cite related papers using this BibTeX:
@article{he2026taxonomy,
title={Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models},
author={He, Hulingxiao and Tan, Zhi and Peng, Yuxin},
journal={arXiv preprint arXiv:2603.00431},
year={2026}
}
@article{he2026fine,
title={Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning},
author={He, Hulingxiao and Geng, Zijun and Peng, Yuxin},
journal={arXiv preprint arXiv:2602.07605},
year={2026}
}
@article{he2025analyzing,
title={Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models},
author={He, Hulingxiao and Li, Geng and Geng, Zijun and Xu, Jinglin and Peng, Yuxin},
journal={arXiv preprint arXiv:2501.15140},
year={2025}
}
This project is licensed under the MIT License.