StevenHH2000/Fine-R1-Stage2-data

Dataset

This is the official release of the training dataset of Stage 2 (TAPO) for paper Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning. Code is available at https://github.com/PKU-ICST-MIPL/FineR1_ICLR2026.

1

3 commits

1 linked in READMEs

updated Feb 11, 2026

See the code

README

This is the official release of the training dataset of Stage 2 (TAPO) for paper Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning. Code is available at https://github.com/PKU-ICST-MIPL/FineR1_ICLR2026.

Data Source

Training

  • We randomly sample 4-shot data per seen category from the following 6 datasets to construct our TAPO training dataset.
    • FGVC-Aircraft ✈️
    • CaltechUCSD Bird-200 🦤
    • Stanford Car-196 🚙
    • Stanford Dog-120 🐕
    • Flower-102 🌼
    • Oxford-IIIT Pet-37 🐈

Data Fields

  • images: input image(s)
    • data type: List
  • problem: input question or statement
    • data type: String
  • answer: ground-truth answer
    • data type: String
  • image_noisy: positive image(s) from the same sub-category
    • data type: List
  • image_aug: negative image(s) from a different sub-category
    • data type: List
multimodal
reinforcement-learning
vision

StevenHH2000/Fine-R1-Stage2-data

Dataset

This is the official release of the training dataset of Stage 2 (TAPO) for paper Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning. Code is available at https://github.com/PKU-ICST-MIPL/FineR1_ICLR2026.

1

3 commits

1 linked in READMEs

updated Feb 11, 2026

See the code

README

This is the official release of the training dataset of Stage 2 (TAPO) for paper Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning. Code is available at https://github.com/PKU-ICST-MIPL/FineR1_ICLR2026.

Data Source

Training

  • We randomly sample 4-shot data per seen category from the following 6 datasets to construct our TAPO training dataset.
    • FGVC-Aircraft ✈️
    • CaltechUCSD Bird-200 🦤
    • Stanford Car-196 🚙
    • Stanford Dog-120 🐕
    • Flower-102 🌼
    • Oxford-IIIT Pet-37 🐈

Data Fields

  • images: input image(s)
    • data type: List
  • problem: input question or statement
    • data type: String
  • answer: ground-truth answer
    • data type: String
  • image_noisy: positive image(s) from the same sub-category
    • data type: List
  • image_aug: negative image(s) from a different sub-category
    • data type: List
multimodal
reinforcement-learning
vision