Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
Python
278
13 commits
updated Sep 22, 2026

conda create -n geo-vista python==3.10 -y
conda activate geo-vista
bash setup.sh
We use Tavily during inference and training. You can sign up for a free account and get your Tavily API key, and then update the TAVILY_API_KEY variable of the .env file. You can run bash examples/search_test.sh to verify your API key is working.
Download from HuggingFace, place it in the ./.temp/checkpoints/LibraTree/GeoVista-RL-6k-7B directory.
python3 scripts/download_hf.py \
--model LibraTree/GeoVista-RL-6k-7B \
--local_model_dir .temp/checkpoints/
then deploy the GeoVista model with vllm:
bash inference/vllm_deploy_geovista_rl_6k.sh

export VLLM_PORT=8000
export VLLM_HOST="localhost"
# apply env variables
set -a; source .env; set +a;
python examples/infer_example.py \
--multimodal_input examples/geobench-example.png \
--question "Please analyze where is the place."
You will see the model's thinking trajectory and final answer in the console output.

GeoBench is the first high-resolution, multi-source, globally annotated dataset to evaluate agentic models’ general geolocalization ability.
| Benchmark | Year | GC | RC | HR | DV | NE |
|---|---|---|---|---|---|---|
| Im2GPS | 2008 | ✓ | ||||
| YFCC4k | 2017 | ✓ | ||||
| Google Landmarks v2 | 2020 | ✓ | ||||
| VIGOR | 2022 | ✓ | ||||
| OSV-5M | 2024 | ✓ | ✓ | ✓ | ||
| GeoComp | 2025 | ✓ | ✓ | ✓ | ||
| GeoBench (ours) | 2025 | ✓ | ✓ | ✓ | ✓ | ✓ |
We provide the whole inference and evaluation pipeline for GeoVista on GeoBench.
./.temp/datasets directory.python3 scripts/download_hf.py \
--dataset LibraTree/GeoVistaBench \
--local_dataset_dir ./.temp/datasets
./.temp/checkpoints directory.python3 scripts/download_hf.py \
--model LibraTree/GeoVista-RL-12k-7B \
--local_model_dir .temp/checkpoints/
bash inference/vllm_deploy.sh
bash inference/run_inference.sh
After running the above commands, you should be able to see the inference results in the specified output directory, e.g., ./.temp/outputs/geobench/geovista-rl-12k-7b/, which contains the inference_<timestamp>.jsonl file with the inference results.
MODEL_NAME=geovista-rl-12k-7b
BENCHMARK=geobench
EVALUATION_RESULT=".temp/outputs/${BENCHMARK}/${MODEL_NAME}/evaluation.jsonl"
python3 eval/eval_infer_geolocation.py \
--pred_jsonl <The inference file path> \
--out_jsonl ${EVALUATION_RESULT}\
--dataset_dir .temp/datasets/${BENCHMARK} \
--num_samples 1500 \
--model_verifier \
--no_eval_accurate_dist \
--timeout 120 --debug | tee .temp/outputs/${BENCHMARK}/${MODEL_NAME}/evaluation.log 2>&1
You can acclerate the evaluation process by changing the workers argument in the above command (default is 1):
--workers 8 \
GeoVista is trained in two stages: (1) Cold‑Start supervised fine‑tuning (SFT) and (2) Reinforcement Learning.
scripts/sft.py.bash scripts/sft.sh
Please consider citing our paper and starring this repo if you find them helpful. Thank you!
@misc{wang2025geovistawebaugmentedagenticvisual,
title={GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization},
author={Yikun Wang and Zuyan Liu and Ziyi Wang and Han Hu and Pengfei Liu and Yongming Rao},
year={2025},
eprint={2511.15705},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.15705},
}
Python
98.6%
Shell
1.4%
Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
Python
278
13 commits
updated Sep 22, 2026

conda create -n geo-vista python==3.10 -y
conda activate geo-vista
bash setup.sh
We use Tavily during inference and training. You can sign up for a free account and get your Tavily API key, and then update the TAVILY_API_KEY variable of the .env file. You can run bash examples/search_test.sh to verify your API key is working.
Download from HuggingFace, place it in the ./.temp/checkpoints/LibraTree/GeoVista-RL-6k-7B directory.
python3 scripts/download_hf.py \
--model LibraTree/GeoVista-RL-6k-7B \
--local_model_dir .temp/checkpoints/
then deploy the GeoVista model with vllm:
bash inference/vllm_deploy_geovista_rl_6k.sh

export VLLM_PORT=8000
export VLLM_HOST="localhost"
# apply env variables
set -a; source .env; set +a;
python examples/infer_example.py \
--multimodal_input examples/geobench-example.png \
--question "Please analyze where is the place."
You will see the model's thinking trajectory and final answer in the console output.

GeoBench is the first high-resolution, multi-source, globally annotated dataset to evaluate agentic models’ general geolocalization ability.
| Benchmark | Year | GC | RC | HR | DV | NE |
|---|---|---|---|---|---|---|
| Im2GPS | 2008 | ✓ | ||||
| YFCC4k | 2017 | ✓ | ||||
| Google Landmarks v2 | 2020 | ✓ | ||||
| VIGOR | 2022 | ✓ | ||||
| OSV-5M | 2024 | ✓ | ✓ | ✓ | ||
| GeoComp | 2025 | ✓ | ✓ | ✓ | ||
| GeoBench (ours) | 2025 | ✓ | ✓ | ✓ | ✓ | ✓ |
We provide the whole inference and evaluation pipeline for GeoVista on GeoBench.
./.temp/datasets directory.python3 scripts/download_hf.py \
--dataset LibraTree/GeoVistaBench \
--local_dataset_dir ./.temp/datasets
./.temp/checkpoints directory.python3 scripts/download_hf.py \
--model LibraTree/GeoVista-RL-12k-7B \
--local_model_dir .temp/checkpoints/
bash inference/vllm_deploy.sh
bash inference/run_inference.sh
After running the above commands, you should be able to see the inference results in the specified output directory, e.g., ./.temp/outputs/geobench/geovista-rl-12k-7b/, which contains the inference_<timestamp>.jsonl file with the inference results.
MODEL_NAME=geovista-rl-12k-7b
BENCHMARK=geobench
EVALUATION_RESULT=".temp/outputs/${BENCHMARK}/${MODEL_NAME}/evaluation.jsonl"
python3 eval/eval_infer_geolocation.py \
--pred_jsonl <The inference file path> \
--out_jsonl ${EVALUATION_RESULT}\
--dataset_dir .temp/datasets/${BENCHMARK} \
--num_samples 1500 \
--model_verifier \
--no_eval_accurate_dist \
--timeout 120 --debug | tee .temp/outputs/${BENCHMARK}/${MODEL_NAME}/evaluation.log 2>&1
You can acclerate the evaluation process by changing the workers argument in the above command (default is 1):
--workers 8 \
GeoVista is trained in two stages: (1) Cold‑Start supervised fine‑tuning (SFT) and (2) Reinforcement Learning.
scripts/sft.py.bash scripts/sft.sh
Please consider citing our paper and starring this repo if you find them helpful. Thank you!
@misc{wang2025geovistawebaugmentedagenticvisual,
title={GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization},
author={Yikun Wang and Zuyan Liu and Ziyi Wang and Han Hu and Pengfei Liu and Yongming Rao},
year={2025},
eprint={2511.15705},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.15705},
}
Python
98.6%
Shell
1.4%