bfshi/llava-v1.5-13b-s2-lora

Model

2

stars

3

commits

2

linked in READMEs

May 3, 2024

updated

endpoints_compatible
llava
text-generation
transformers

README

CODE

When Do We Not Need Larger Vision Models?

Model

This is a LLaVA-v1.5-13b model trained with S2-Wrapper, a simple approach to enable any vision model to perceive high-resolution images. We use image resolutions of up to 1008x1008 for this model.

Training

The training pipeline and dataset completely follow LLaVA-v1.5. We use LoRA to fine-tune the model.

Benchmarking

VersionSizeScheduleCheckpointVQAv2VizWizTextVQAMMMU-valMathVistaMM-BenchSEEDMM-Vet
LLaVA-1.513Bfull_ft-1eliuhaotian/llava-v1.5-13b80.053.661.336.427.667.768.236.1
LLaVA-1.513Blora-1eliuhaotian/llava-v1.5-13b-lora80.058.960.2--68.5-38.3
LLaVA-1.5-S213Blora-1ethis model80.956.063.137.427.867.968.936.4

License

Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

Contributors

bfshi

3 commits

bfshi/llava-v1.5-13b-s2-lora

Model

2

stars

3

commits

2

linked in READMEs

May 3, 2024

updated

endpoints_compatible
llava
text-generation
transformers

README

CODE

When Do We Not Need Larger Vision Models?

Model

This is a LLaVA-v1.5-13b model trained with S2-Wrapper, a simple approach to enable any vision model to perceive high-resolution images. We use image resolutions of up to 1008x1008 for this model.

Training

The training pipeline and dataset completely follow LLaVA-v1.5. We use LoRA to fine-tune the model.

Benchmarking

VersionSizeScheduleCheckpointVQAv2VizWizTextVQAMMMU-valMathVistaMM-BenchSEEDMM-Vet
LLaVA-1.513Bfull_ft-1eliuhaotian/llava-v1.5-13b80.053.661.336.427.667.768.236.1
LLaVA-1.513Blora-1eliuhaotian/llava-v1.5-13b-lora80.058.960.2--68.5-38.3
LLaVA-1.5-S213Blora-1ethis model80.956.063.137.427.867.968.936.4

License

Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

Contributors

bfshi

3 commits