StevenHH2000/iNat21-1shot-fewshots

Dataset

This is the official release of the training dataset for paper Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models. Code is available at https://github.com/PKU-ICST-MIPL/TARA_CVPR2026.

1

2 commits

1 linked in READMEs

updated Mar 19, 2026

See the code

README

This is the official release of the training dataset for paper Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models. Code is available at https://github.com/PKU-ICST-MIPL/TARA_CVPR2026.

Data Source

Training

  • We randomly sample 1-shot data per category from iNaturalist2021 dataset.

Data Fields

  • image: input image(s)
    • data type: dict
  • problem: input question
    • data type: string
  • label: coarse-to-fine categories
    • data type: string
  • image_width: image width
    • data type: int64
  • image_height: image height
    • data type: int64
multimodal
vision

StevenHH2000/iNat21-1shot-fewshots

Dataset

This is the official release of the training dataset for paper Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models. Code is available at https://github.com/PKU-ICST-MIPL/TARA_CVPR2026.

1

2 commits

1 linked in READMEs

updated Mar 19, 2026

See the code

README

This is the official release of the training dataset for paper Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models. Code is available at https://github.com/PKU-ICST-MIPL/TARA_CVPR2026.

Data Source

Training

  • We randomly sample 1-shot data per category from iNaturalist2021 dataset.

Data Fields

  • image: input image(s)
    • data type: dict
  • problem: input question
    • data type: string
  • label: coarse-to-fine categories
    • data type: string
  • image_width: image width
    • data type: int64
  • image_height: image height
    • data type: int64
multimodal
vision