[ACL 2026 Oral] UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
176
stars
4
commits
Python
primary language
Jun 24, 2026
updated
UniversalRAG is a novel any-to-any RAG framework that retrieves across multiple modalities and granularities by introducing a modality-aware routing mechanism that dynamically identifies the most appropriate modality-specific corpus for each query, effectively addressing the limitations posed by modality gaps and fixed-granularity retrieval.
To set up the environment, we recommend using uv for a fast and deterministic setup. All dependencies are specified in pyproject.toml and pinned in uv.lock.
git clone https://github.com/wgcyeo/UniversalRAG.git
cd UniversalRAG
uv sync
source .venv/bin/activate
bash script/0_dataset.sh
Run the following command to extract embeddings for all queries and corpora across diverse modalities:
bash script/1_preprocess.sh
Set the CUDA_DEVICES variable (e.g., CUDA_DEVICES="0 1 2 3") to match the GPUs available on your system.
To train a training-based router (e.g., Qwen3-VL-2B-Instruct or T5Gemma 2 270M), run:
bash script/2_train.sh {model-name}
Replace {model-name} with the desired router model. Supported options are qwen3-vl-2b, internvl3_5-1b, and t5gemma-270m.
[!WARNING] To use T5Gemma 2, install Transformers v5 or later:
uv pip install "transformers>=5"
To route queries with either a training-free router (e.g., GPT-5) or a training-based router, run:
bash script/3_route.sh {model-name}
Replace {model-name} with the router model to use. The supported training-free option is gpt-5; training-based options are qwen3-vl-2b, internvl3_5-1b, and t5gemma-270m.
To generate results for routed queries, run the following script:
bash script/4_eval.sh \
--model-path {model-path} \
--router-model {router-model} \
--target {target}
{model-path}: Path or identifier of the LVLM model to use (e.g., Qwen/Qwen3-VL-8B-Instruct).{router-model}: Router model name used in the routing stage (e.g., gpt-5, qwen3-vl-2b, internvl3_5-1b, or t5gemma-270m).{target}: Target dataset for evaluation (e.g., mmlu).Use bash script/4_eval.sh -h to see all available options and descriptions.
Example:
bash script/4_eval.sh \
--model-path Qwen/Qwen3-VL-8B-Instruct \
--router-model qwen3-vl-2b \
--target mmlu
If you find this work useful, please consider citing our paper:
@inproceedings{yeo2026universalrag,
title = {UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities},
author = {Yeo, Woongyeong and Kim, Kangsan and Jeong, Soyeong and Baek, Jinheon and Hwang, Sung Ju},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
month = {July},
year = {2026},
pages = {3843-3871}
}
4 commits
Python
98.4%
Shell
1.6%
[ACL 2026 Oral] UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
176
stars
4
commits
Python
primary language
Jun 24, 2026
updated
UniversalRAG is a novel any-to-any RAG framework that retrieves across multiple modalities and granularities by introducing a modality-aware routing mechanism that dynamically identifies the most appropriate modality-specific corpus for each query, effectively addressing the limitations posed by modality gaps and fixed-granularity retrieval.
To set up the environment, we recommend using uv for a fast and deterministic setup. All dependencies are specified in pyproject.toml and pinned in uv.lock.
git clone https://github.com/wgcyeo/UniversalRAG.git
cd UniversalRAG
uv sync
source .venv/bin/activate
bash script/0_dataset.sh
Run the following command to extract embeddings for all queries and corpora across diverse modalities:
bash script/1_preprocess.sh
Set the CUDA_DEVICES variable (e.g., CUDA_DEVICES="0 1 2 3") to match the GPUs available on your system.
To train a training-based router (e.g., Qwen3-VL-2B-Instruct or T5Gemma 2 270M), run:
bash script/2_train.sh {model-name}
Replace {model-name} with the desired router model. Supported options are qwen3-vl-2b, internvl3_5-1b, and t5gemma-270m.
[!WARNING] To use T5Gemma 2, install Transformers v5 or later:
uv pip install "transformers>=5"
To route queries with either a training-free router (e.g., GPT-5) or a training-based router, run:
bash script/3_route.sh {model-name}
Replace {model-name} with the router model to use. The supported training-free option is gpt-5; training-based options are qwen3-vl-2b, internvl3_5-1b, and t5gemma-270m.
To generate results for routed queries, run the following script:
bash script/4_eval.sh \
--model-path {model-path} \
--router-model {router-model} \
--target {target}
{model-path}: Path or identifier of the LVLM model to use (e.g., Qwen/Qwen3-VL-8B-Instruct).{router-model}: Router model name used in the routing stage (e.g., gpt-5, qwen3-vl-2b, internvl3_5-1b, or t5gemma-270m).{target}: Target dataset for evaluation (e.g., mmlu).Use bash script/4_eval.sh -h to see all available options and descriptions.
Example:
bash script/4_eval.sh \
--model-path Qwen/Qwen3-VL-8B-Instruct \
--router-model qwen3-vl-2b \
--target mmlu
If you find this work useful, please consider citing our paper:
@inproceedings{yeo2026universalrag,
title = {UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities},
author = {Yeo, Woongyeong and Kim, Kangsan and Jeong, Soyeong and Baek, Jinheon and Hwang, Sung Ju},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
month = {July},
year = {2026},
pages = {3843-3871}
}
4 commits
Python
98.4%
Shell
1.6%