[TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
See the codeTPAMI 2026
Ego-R1 is a comprehensive research framework that combines reinforcement learning-based tool-use reasoning with egocentric video analysis capabilities.
This repository provides:
Ego-R1/
├── cott_gen/ # Chain-of-Tool-Thought generation for egocentric video QA
│ ├── main.py # Main agent runner with multi-turn reasoning
│ ├── tools.py # Tool implementations (RAG, Video-LLM, VLM)
│ ├── utils.py # Utility functions and data processing
│ ├── prompts.py # System and reasoning prompts
│ ├── postprocess.py # Data postprocessing and analysis
│ └── environment.yml # Conda environment for autogen
├── LLaMA-Factory/ # LLM fine-tuning framework (submodule)
├── Ego-R1-Agent/ # RL framework for reasoning + search LLMs
│ ├── train_grpo.sh # GRPO training script
│ ├── train_ppo.sh # PPO training script
│ ├── eval/ # Inference and evaluation scripts
│ └── verl/ # veRL framework components
├── data/ # Ego-R1 dataset (should be downloaded from HF)
│ ├── Ego-CoTT-25K/ # 25K Chain-of-Tool-Thought for SFT
│ ├── Ego-QA-4.4K/ # 4.4K QA pairs for RL training
│ └── Ego-CoTT-raw/ # Raw data in multiple formats
├── scripts/ # Training and generation scripts
│ ├── train/ # SFT training scripts
│ └── gen/ # Data generation scripts
└── api/ # API components for RAG and visual tools
├── rag/ # RAG-related API components
└── visual_tools/ # Multi-modal visual tool APIs
huggingface-cli download Ego-R1/Ego-R1-Data --local-dir data --repo-type dataset
i. Set Environment
cd api/rag
pip install -e .
Make sure to install FFmpeg beforehand, as it is required for the visual tools to function properly.
ii. Prepare the Data For Egoschema and Videomme benchmark
huggingface-cli download Ego-R1/h-rag_database --local-dir data --repo-type dataset
Unzip the Videomme and Egoschema videos.
iii. Setup API
Set GPT Key
export AZURE_OPENAI_ENDPOINT=ENDPOINT
export AZURE_OPENAI_API_KEY=KEY
Start RAG
For Egolife/Ego-R1:
rag/configs/egolife.yaml:
base:
data_dir: data/egolife # set to h-rag_database/egolife
python api_for_egolife.py
For Egoschema:
python api_for_egoschema.py --min_log_dir=h-rag_database/egoschema --port 6001 # default
For Videomme:
python api_for_videomme.py --min_log_dir=h-rag_database/videomme/videomme_10min --sec_log_dir=h-rag_database/videomme/videomme_30s --port 7001 # default
iv. Start Visual API
Set Config
visual_tools/configs.yaml for EgoLife, Egoschema, and Videomme videos separately:
data_dir: "/path/to/egolife"
data_dir: "/path/to/videomme"
data_dir: "/path/to/egoschema"
gemini_api_keys: ["your-gemini-api-key-1", "your-gemini-api-key-2"]
Run API
python api.py
python xxxx_videollm_llava/llava_video.py
# One-line installation
cd cott_gen
conda env create -f environment.yml
conda activate autogen
# Or install step by step:
# conda create -n autogen python=3.10
# conda activate autogen
# pip install -U autogenstudio==0.6.1
# pip install future google-genai
cd LLaMA-Factory
pip install -e ".[torch,metrics]"
conda create -n egor1 python=3.9
conda activate egor1
# Install PyTorch (optional - vllm can handle this)
pip install torch==2.4.0 --index-url https://download.pytorch.org/whl/cu121
# verl
pip install -e .
# flash attention 2
pip3 install flash-attn --no-build-isolation
pip install wandb google-genai
You can follow Search-R1 to build the environment as well.
bash Ego-R1-Agent/utils/serve.sh
conda activate egor1
# with a summary model
bash Ego-R1-Agent/eval/infer_bench_summ.sh
# or you can go with a basic one
# python infer.py --arg1 xxx --arg2 xxx
# Prepare data
mkdir -p LLaMA-Factory/data
cp data/Ego-CoTT-25K/train-cott.json LLaMA-Factory/data/
# Train model
conda activate llamafactory
cd LLaMA-Factory
llamafactory-cli train examples/train_full/qwen.yaml
# Prepare data
mkdir -p Ego-R1-Agent/data
cp data/Ego-CoTT-raw/*.parquet Ego-R1-Agent/data/
# Start RL training
conda activate egor1
cd Ego-R1-Agent
bash train_grpo.sh # For GRPO training
# Generate reasoning traces with multi-modal tools
conda activate autogen
bash scripts/gen/run_data_gen.sh
The Ego-R1 agent uses a structured chain-of-tool-thought approach:
{
"name": "rag",
"arguments": {
"level": "day", # or "week", "hour"
"keywords": ["cooking", "kitchen"],
"start_time": "DAY1_11210217",
"query_time": "DAY1_11220217"
}
}
{
"name": "video_llm",
"arguments": {
"question": "What cooking action is being performed?",
"range": "DAY1_11210217-DAY1_11220217"
}
}
{
"name": "vlm",
"arguments": {
"question": "What objects are visible on the table?",
"timestamp": "DAY1_11210217"
}
}
This project builds upon several excellent open-source frameworks:
This project is licensed under the Apache License 2.0. See the LICENSE files in individual components for details.
Contributions are welcome! Please feel free to submit issues, feature requests, or pull requests to help improve this research framework.
If you have any queries, feel free to contact: Shulin Tian (shulin002@ntu.edu.sg) & Ruiqi Wang (rwa135@sfu.ca)
@article{tian2026ego,
title={Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning},
author={Tian, Shulin and Wang, Ruiqi and Guo, Hongming and Wu, Penghao and Dong, Yuhao and Wang, Xiuying and Yang, Jingkang and Zhang, Hao and Zhu, Hongyuan and Liu, Ziwei},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2026},
publisher={IEEE}
}
[TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
See the codeTPAMI 2026
Ego-R1 is a comprehensive research framework that combines reinforcement learning-based tool-use reasoning with egocentric video analysis capabilities.
This repository provides:
Ego-R1/
├── cott_gen/ # Chain-of-Tool-Thought generation for egocentric video QA
│ ├── main.py # Main agent runner with multi-turn reasoning
│ ├── tools.py # Tool implementations (RAG, Video-LLM, VLM)
│ ├── utils.py # Utility functions and data processing
│ ├── prompts.py # System and reasoning prompts
│ ├── postprocess.py # Data postprocessing and analysis
│ └── environment.yml # Conda environment for autogen
├── LLaMA-Factory/ # LLM fine-tuning framework (submodule)
├── Ego-R1-Agent/ # RL framework for reasoning + search LLMs
│ ├── train_grpo.sh # GRPO training script
│ ├── train_ppo.sh # PPO training script
│ ├── eval/ # Inference and evaluation scripts
│ └── verl/ # veRL framework components
├── data/ # Ego-R1 dataset (should be downloaded from HF)
│ ├── Ego-CoTT-25K/ # 25K Chain-of-Tool-Thought for SFT
│ ├── Ego-QA-4.4K/ # 4.4K QA pairs for RL training
│ └── Ego-CoTT-raw/ # Raw data in multiple formats
├── scripts/ # Training and generation scripts
│ ├── train/ # SFT training scripts
│ └── gen/ # Data generation scripts
└── api/ # API components for RAG and visual tools
├── rag/ # RAG-related API components
└── visual_tools/ # Multi-modal visual tool APIs
huggingface-cli download Ego-R1/Ego-R1-Data --local-dir data --repo-type dataset
i. Set Environment
cd api/rag
pip install -e .
Make sure to install FFmpeg beforehand, as it is required for the visual tools to function properly.
ii. Prepare the Data For Egoschema and Videomme benchmark
huggingface-cli download Ego-R1/h-rag_database --local-dir data --repo-type dataset
Unzip the Videomme and Egoschema videos.
iii. Setup API
Set GPT Key
export AZURE_OPENAI_ENDPOINT=ENDPOINT
export AZURE_OPENAI_API_KEY=KEY
Start RAG
For Egolife/Ego-R1:
rag/configs/egolife.yaml:
base:
data_dir: data/egolife # set to h-rag_database/egolife
python api_for_egolife.py
For Egoschema:
python api_for_egoschema.py --min_log_dir=h-rag_database/egoschema --port 6001 # default
For Videomme:
python api_for_videomme.py --min_log_dir=h-rag_database/videomme/videomme_10min --sec_log_dir=h-rag_database/videomme/videomme_30s --port 7001 # default
iv. Start Visual API
Set Config
visual_tools/configs.yaml for EgoLife, Egoschema, and Videomme videos separately:
data_dir: "/path/to/egolife"
data_dir: "/path/to/videomme"
data_dir: "/path/to/egoschema"
gemini_api_keys: ["your-gemini-api-key-1", "your-gemini-api-key-2"]
Run API
python api.py
python xxxx_videollm_llava/llava_video.py
# One-line installation
cd cott_gen
conda env create -f environment.yml
conda activate autogen
# Or install step by step:
# conda create -n autogen python=3.10
# conda activate autogen
# pip install -U autogenstudio==0.6.1
# pip install future google-genai
cd LLaMA-Factory
pip install -e ".[torch,metrics]"
conda create -n egor1 python=3.9
conda activate egor1
# Install PyTorch (optional - vllm can handle this)
pip install torch==2.4.0 --index-url https://download.pytorch.org/whl/cu121
# verl
pip install -e .
# flash attention 2
pip3 install flash-attn --no-build-isolation
pip install wandb google-genai
You can follow Search-R1 to build the environment as well.
bash Ego-R1-Agent/utils/serve.sh
conda activate egor1
# with a summary model
bash Ego-R1-Agent/eval/infer_bench_summ.sh
# or you can go with a basic one
# python infer.py --arg1 xxx --arg2 xxx
# Prepare data
mkdir -p LLaMA-Factory/data
cp data/Ego-CoTT-25K/train-cott.json LLaMA-Factory/data/
# Train model
conda activate llamafactory
cd LLaMA-Factory
llamafactory-cli train examples/train_full/qwen.yaml
# Prepare data
mkdir -p Ego-R1-Agent/data
cp data/Ego-CoTT-raw/*.parquet Ego-R1-Agent/data/
# Start RL training
conda activate egor1
cd Ego-R1-Agent
bash train_grpo.sh # For GRPO training
# Generate reasoning traces with multi-modal tools
conda activate autogen
bash scripts/gen/run_data_gen.sh
The Ego-R1 agent uses a structured chain-of-tool-thought approach:
{
"name": "rag",
"arguments": {
"level": "day", # or "week", "hour"
"keywords": ["cooking", "kitchen"],
"start_time": "DAY1_11210217",
"query_time": "DAY1_11220217"
}
}
{
"name": "video_llm",
"arguments": {
"question": "What cooking action is being performed?",
"range": "DAY1_11210217-DAY1_11220217"
}
}
{
"name": "vlm",
"arguments": {
"question": "What objects are visible on the table?",
"timestamp": "DAY1_11210217"
}
}
This project builds upon several excellent open-source frameworks:
This project is licensed under the Apache License 2.0. See the LICENSE files in individual components for details.
Contributions are welcome! Please feel free to submit issues, feature requests, or pull requests to help improve this research framework.
If you have any queries, feel free to contact: Shulin Tian (shulin002@ntu.edu.sg) & Ruiqi Wang (rwa135@sfu.ca)
@article{tian2026ego,
title={Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning},
author={Tian, Shulin and Wang, Ruiqi and Guo, Hongming and Wu, Penghao and Dong, Yuhao and Wang, Xiuying and Yang, Jingkang and Zhang, Hao and Zhu, Hongyuan and Liu, Ziwei},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2026},
publisher={IEEE}
}