[ECCV 2024] FlexAttention for Efficient High-Resolution Vision-Language Models
Python
51
12 commits
updated Jan 8, 2025
[Project Page] [Paper]

This repository contains the official code for FlexAttention for Efficient High-Resolution Vision-Language Models.
conda create -n flexattention python=3.9
conda activate flexattention
pip install -e .
pip install -e ".[train]"
pip install -e ./transformers
You can download our 7B model checkpoint from huggingface and put it into checkpoints folder.
datasets/eval/textvqa.torchrun --nproc_per_node 3 scripts/evaluation/eval_textvqa.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn
It will generate a file similar to answer_textvqa_llava-v1.5-7b-flexattn_xxx.jsonl on the folder root.
bash scripts/evaluation/get_textvqa_score.sh ANSWER_FILE
git lfs install
git clone https://huggingface.co/datasets/craigwu/vstar_bench
# Attribute
torchrun --nproc_per_node 3 scripts/evaluation/eval_vbench.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn --subset direct_attributes
# Spatial
torchrun --nproc_per_node 3 scripts/evaluation/eval_vbench.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn --subset relative_position
Download the dataset from here, and extract it to datasets/eval/.
Run the multi-gpu inference:
torchrun --nproc_per_node 3 scripts/evaluation/eval_magnifier.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn
First, follow LLaVA's instruction to prepare the image and annotation. The overall folder structure will look like this:
playground
├── llava_v1_5_mix665k
├── llava_v1_5_mix665k.json
├── coco
│ └── train2017
├── gqa
│ └── images
├── ocr_vqa
│ └── images
├── textvqa
│ └── train_images
└── vg
├── VG_100K
└── VG_100K_2
Then, run the data cleaning script to clean the data.
python tools/prepare_data.py
In this script, we perform the following actions:
<image> tag for samples containing only text.You can directly download the prepared file here.
Finally, run the training script:
bash scripts/train/llava-v1.5-7b-flexattn.sh
LLaVA: the codebase that our project build on. Thanks for their amazing code and model.
If our work is useful or relevant to your research, please kindly recognize our contributions by citing our paper:
@inproceedings{li2025flexattention,
title={Flexattention for efficient high-resolution vision-language models},
author={Li, Junyan and Chen, Delin and Cai, Tianle and Chen, Peihao and Hong, Yining and Chen, Zhenfang and Shen, Yikang and Gan, Chuang},
booktitle={European Conference on Computer Vision},
pages={286--302},
year={2025},
organization={Springer}
}
Python
99.4%
[ECCV 2024] FlexAttention for Efficient High-Resolution Vision-Language Models
Python
51
12 commits
updated Jan 8, 2025
[Project Page] [Paper]

This repository contains the official code for FlexAttention for Efficient High-Resolution Vision-Language Models.
conda create -n flexattention python=3.9
conda activate flexattention
pip install -e .
pip install -e ".[train]"
pip install -e ./transformers
You can download our 7B model checkpoint from huggingface and put it into checkpoints folder.
datasets/eval/textvqa.torchrun --nproc_per_node 3 scripts/evaluation/eval_textvqa.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn
It will generate a file similar to answer_textvqa_llava-v1.5-7b-flexattn_xxx.jsonl on the folder root.
bash scripts/evaluation/get_textvqa_score.sh ANSWER_FILE
git lfs install
git clone https://huggingface.co/datasets/craigwu/vstar_bench
# Attribute
torchrun --nproc_per_node 3 scripts/evaluation/eval_vbench.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn --subset direct_attributes
# Spatial
torchrun --nproc_per_node 3 scripts/evaluation/eval_vbench.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn --subset relative_position
Download the dataset from here, and extract it to datasets/eval/.
Run the multi-gpu inference:
torchrun --nproc_per_node 3 scripts/evaluation/eval_magnifier.py --dist --model-path checkpoints/llava-v1.5-7b-flexattn --id llava-v1.5-7b-flexattn
First, follow LLaVA's instruction to prepare the image and annotation. The overall folder structure will look like this:
playground
├── llava_v1_5_mix665k
├── llava_v1_5_mix665k.json
├── coco
│ └── train2017
├── gqa
│ └── images
├── ocr_vqa
│ └── images
├── textvqa
│ └── train_images
└── vg
├── VG_100K
└── VG_100K_2
Then, run the data cleaning script to clean the data.
python tools/prepare_data.py
In this script, we perform the following actions:
<image> tag for samples containing only text.You can directly download the prepared file here.
Finally, run the training script:
bash scripts/train/llava-v1.5-7b-flexattn.sh
LLaVA: the codebase that our project build on. Thanks for their amazing code and model.
If our work is useful or relevant to your research, please kindly recognize our contributions by citing our paper:
@inproceedings{li2025flexattention,
title={Flexattention for efficient high-resolution vision-language models},
author={Li, Junyan and Chen, Delin and Cai, Tianle and Chen, Peihao and Hong, Yining and Chen, Zhenfang and Shen, Yikang and Gan, Chuang},
booktitle={European Conference on Computer Vision},
pages={286--302},
year={2025},
organization={Springer}
}
Python
99.4%