BAAI/DenseFusion-1M

Dataset

- [Paper] https://arxiv.org/abs/2407.08303

44

71 commits

2 linked in READMEs

updated Oct 17, 2024

See the code

README

Introduction

  • An image is worth a thousand words". Comprehensive image descriptions are essential for multi-modal perception, while images contains various visual elements of different granularities that are challenging to harness.
  • We propose Perceptural Fusion to integrate the diverse visual perception experts for capturing visual elements and adopt a MLLM as a centric pivot for comprehensive perception.
  • We thereby provide DenseFusion-1M dataset for highly informative image descriptions with various visual details, including rich OCR information, accurate object and position recognition, and external knowledge, etc.

Detaset details: Comprehensive image descriptions obtained through perceptual fusion of different visual experts.

Usage: It is constructed for comprehensive perception ability for multi-modal large language model and show potentials for fine-grained text conditioned image generation.

DenseFusion Dataset Card

  • DenseFusion-1M: DenseFusion-1M/DenseFusion-1M.jsonl is generated by our caption engine through perceptual fusion.
  • DenseFusion-4V-100K: DenseFusion-4V-100k/DenseFusion-4V-100k.jsonl is generated by GPT-4V through perceptual fusion.

This project is under the policy of MIT License.

GPT-4V
MLLM

Contributors

ryanzhangfan

67 commits

FZ
Fan Zhang

3 commits

LX
LI, XIAOTONG

1 commits

BAAI/DenseFusion-1M

Dataset

- [Paper] https://arxiv.org/abs/2407.08303

44

71 commits

2 linked in READMEs

updated Oct 17, 2024

See the code

README

Introduction

  • An image is worth a thousand words". Comprehensive image descriptions are essential for multi-modal perception, while images contains various visual elements of different granularities that are challenging to harness.
  • We propose Perceptural Fusion to integrate the diverse visual perception experts for capturing visual elements and adopt a MLLM as a centric pivot for comprehensive perception.
  • We thereby provide DenseFusion-1M dataset for highly informative image descriptions with various visual details, including rich OCR information, accurate object and position recognition, and external knowledge, etc.

Detaset details: Comprehensive image descriptions obtained through perceptual fusion of different visual experts.

Usage: It is constructed for comprehensive perception ability for multi-modal large language model and show potentials for fine-grained text conditioned image generation.

DenseFusion Dataset Card

  • DenseFusion-1M: DenseFusion-1M/DenseFusion-1M.jsonl is generated by our caption engine through perceptual fusion.
  • DenseFusion-4V-100K: DenseFusion-4V-100k/DenseFusion-4V-100k.jsonl is generated by GPT-4V through perceptual fusion.

This project is under the policy of MIT License.

GPT-4V
MLLM

Contributors

ryanzhangfan

67 commits

FZ
Fan Zhang

3 commits

LX
LI, XIAOTONG

1 commits