bfshi/llava-v1.5-7b-s2-lora

Model

1

stars

3

commits

2

linked in READMEs

May 3, 2024

updated

endpoints_compatible
llava
text-generation
transformers

README

CODE

When Do We Not Need Larger Vision Models?

Model

This is a LLaVA-v1.5-7b model trained with S2-Wrapper, a simple approach to enable any vision model to perceive high-resolution images. We use image resolutions of up to 1008x1008 for this model.

Training

The training pipeline and dataset completely follow LLaVA-v1.5. We use LoRA to fine-tune the model.

Benchmarking

VersionSizeScheduleCheckpointVQAv2VizWizTextVQAMMMU-valMathVistaMM-BenchSEEDMM-Vet
LLaVA-1.57Bfull_ft-1eliuhaotian/llava-v1.5-7b78.550.058.236.225.264.365.731.1
LLaVA-1.57Blora-1eliuhaotian/llava-v1.5-7b-lora79.147.858.2--66.1-30.2
LLaVA-1.5-S27Blora-1ethis model80.050.161.037.725.366.267.932.4

License

Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

Contributors

bfshi

3 commits

bfshi/llava-v1.5-7b-s2-lora

Model

1

stars

3

commits

2

linked in READMEs

May 3, 2024

updated

endpoints_compatible
llava
text-generation
transformers

README

CODE

When Do We Not Need Larger Vision Models?

Model

This is a LLaVA-v1.5-7b model trained with S2-Wrapper, a simple approach to enable any vision model to perceive high-resolution images. We use image resolutions of up to 1008x1008 for this model.

Training

The training pipeline and dataset completely follow LLaVA-v1.5. We use LoRA to fine-tune the model.

Benchmarking

VersionSizeScheduleCheckpointVQAv2VizWizTextVQAMMMU-valMathVistaMM-BenchSEEDMM-Vet
LLaVA-1.57Bfull_ft-1eliuhaotian/llava-v1.5-7b78.550.058.236.225.264.365.731.1
LLaVA-1.57Blora-1eliuhaotian/llava-v1.5-7b-lora79.147.858.2--66.1-30.2
LLaVA-1.5-S27Blora-1ethis model80.050.161.037.725.366.267.932.4

License

Llama 2 is licensed under the LLAMA 2 Community License, Copyright (c) Meta Platforms, Inc. All Rights Reserved.

Contributors

bfshi

3 commits