Official code for TEOChat, the first vision-language assistant for temporal earth observation data (ICLR 2025).
See the code
TEOChat is the first language and vision assistant that can engage in conversation about sequences of temporal earth observation imagery, and exhibits impressive performance on multiple temporal instruction-following tasks.
We introduce a new instruction-following dataset for temporal EO data called TEOChatlas which we use to train TEOChat. TEOChatlas contains 554,071 examples spanning dozens of temporal instruction-following tasks.
We design TEOChat to use a LLaVA-style architecture, combining a temporally shared vision encoder with a LLaMA 2 LLM connected through an MLP vision-language projector
We provide an online demo in Huggingface Spaces.
You can also run the demo locally by running the following command:
python videollava/serve/teochat_demo.py
We demonstrate that TEOChat:
git clone https://github.com/ermongroup/TEOChat.git
cd TEOChat
conda create -n teochat python=3.9 -y
conda activate teochat
pip install --upgrade pip # enable PEP 660 support
pip install -e .
pip install git+https://github.com/facebookresearch/pytorchvideo
The training & validating instructions, including how to download the TEOChatlas dataset, are in TRAIN_AND_VALIDATE.md.
You can use the following code to run inference with TEOChat on GPU:
from videollava.eval.eval import load_model
from videollava.eval.inference import run_inference_single
tokenizer, model, processor = load_model(model_path="jirvin16/TEOChat", model_base=None, load_8bit=True, device='cuda')
# A list of image paths, all are fed into the model.
image_paths = ["videollava/serve/examples/xBD_cls_1.png", "videollava/serve/examples/xBD_cls_2.png"]
# Note you must include the video tag <video> in the input string otherwise the model will not process the images.
inp = "These are two satellite images in chronological order: <video> Classify the level of damage experienced by the building at location [0, 8, 49, 53]."
response = run_inference_single(model, processor, tokenizer, inp, image_paths)
print(response)
If you find our paper and code useful in your research, please consider giving a star :star: and citation :pencil:.
@inproceedings{irvin2024teochat,
title={TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data},
author={Irvin, Jeremy Andrew and Liu, Emily Ruoyu and Chen, Joyce Chuyi and Dormoy, Ines and Kim, Jinyoung and Khanna, Samar and Zheng, Zhuo and Ermon, Stefano},
booktitle={International Conference on Learning Representations},
year={2025}
}
Python
99.5%
Official code for TEOChat, the first vision-language assistant for temporal earth observation data (ICLR 2025).
See the code
TEOChat is the first language and vision assistant that can engage in conversation about sequences of temporal earth observation imagery, and exhibits impressive performance on multiple temporal instruction-following tasks.
We introduce a new instruction-following dataset for temporal EO data called TEOChatlas which we use to train TEOChat. TEOChatlas contains 554,071 examples spanning dozens of temporal instruction-following tasks.
We design TEOChat to use a LLaVA-style architecture, combining a temporally shared vision encoder with a LLaMA 2 LLM connected through an MLP vision-language projector
We provide an online demo in Huggingface Spaces.
You can also run the demo locally by running the following command:
python videollava/serve/teochat_demo.py
We demonstrate that TEOChat:
git clone https://github.com/ermongroup/TEOChat.git
cd TEOChat
conda create -n teochat python=3.9 -y
conda activate teochat
pip install --upgrade pip # enable PEP 660 support
pip install -e .
pip install git+https://github.com/facebookresearch/pytorchvideo
The training & validating instructions, including how to download the TEOChatlas dataset, are in TRAIN_AND_VALIDATE.md.
You can use the following code to run inference with TEOChat on GPU:
from videollava.eval.eval import load_model
from videollava.eval.inference import run_inference_single
tokenizer, model, processor = load_model(model_path="jirvin16/TEOChat", model_base=None, load_8bit=True, device='cuda')
# A list of image paths, all are fed into the model.
image_paths = ["videollava/serve/examples/xBD_cls_1.png", "videollava/serve/examples/xBD_cls_2.png"]
# Note you must include the video tag <video> in the input string otherwise the model will not process the images.
inp = "These are two satellite images in chronological order: <video> Classify the level of damage experienced by the building at location [0, 8, 49, 53]."
response = run_inference_single(model, processor, tokenizer, inp, image_paths)
print(response)
If you find our paper and code useful in your research, please consider giving a star :star: and citation :pencil:.
@inproceedings{irvin2024teochat,
title={TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data},
author={Irvin, Jeremy Andrew and Liu, Emily Ruoyu and Chen, Joyce Chuyi and Dormoy, Ines and Kim, Jinyoung and Khanna, Samar and Zheng, Zhuo and Ermon, Stefano},
booktitle={International Conference on Learning Representations},
year={2025}
}
Python
99.5%