Describe Anything: Detailed Localized Image and Video Captioning
59
33 commits
2 linked in READMEs
updated Apr 24, 2025
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These tar files can be loaded as a webdataset. Alternatively, you can decompress the tar files and use the json file to load the images without using webdatasets.
This dataset collection includes annotations and images from the following datasets:
Each dataset provides localized descriptions used in the training of Describe Anything Models (DAM).
This dataset is intended to demonstrate and facilitate the understanding and usage of the describe anything models. It should primarily be used for research purposes.
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report security vulnerabilities or NVIDIA AI Concerns here.
Describe Anything: Detailed Localized Image and Video Captioning
59
33 commits
2 linked in READMEs
updated Apr 24, 2025
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These tar files can be loaded as a webdataset. Alternatively, you can decompress the tar files and use the json file to load the images without using webdatasets.
This dataset collection includes annotations and images from the following datasets:
Each dataset provides localized descriptions used in the training of Describe Anything Models (DAM).
This dataset is intended to demonstrate and facilitate the understanding and usage of the describe anything models. It should primarily be used for research purposes.
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report security vulnerabilities or NVIDIA AI Concerns here.