This repository contains the official Pytorch implementation of the paper "Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting".
According to our paper, we conducted two tasks with the following datasets.
There are two options for pre-processing the datasets.
We organize the datasets as follows:
├── datasets
│ | recon
│ ├── DAVIS
│ ├── JPEGImages
│ ├── 480p
│ ├── blackswan
│ ├── blackswan_pts_camera_from_deva
│ ├── ...
│ | edit
│ ├── DAVIS
│ ├── 480p_frames
│ ├── bmx-rider
│ ├── bmx-rider_pts_camera_from_deva
Setting up environments for training contains three parts:
git clone https://github.com/dlsrbgg33/Video-3DGS.git --recursive
cd Video-3DGS
conda create -n video_3dgs python=3.8
conda activate video_3dgs
# install pytorch
pip install torch==1.12.1+cu113 torchvision==0.13.1+cu113 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu113
# install packages & dependencies
bash requirement.sh
Setting up environments for evaluation contains two parts:
cd models/optical_flow/RAFT
bash download_models.sh
cd models/clipscore
git lfs install
git clone https://huggingface.co/openai/clip-vit-large-patch14
git clone https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K
bash sh_recon/davis.sh
To effectively obtain reprentation for video editing, we utilize all the training images for each video scene in this stage.
Arguments:
https://github.com/user-attachments/assets/8eb8e201-ef3b-461c-985b-72d3fa19cd54?width=100&height=100
bash sh_edit/{initial_editor}/{dataset}.sh
We currently support three "initial editors": Text2Video-Zero / TokenFlow / RAVE
We recommend user to install related packages and modules of above initial editors in Video-3DGS framework to conduct initial video editing.
For running TokenFlow efficiently (e.g., edit long video), we borrowed the some strategies from here.
https://github.com/user-attachments/assets/923dec4a-fb23-4c02-a187-a51cb57501a8?width=100&height=100
bash sh_edit/{initial_editor}/davis_re.sh
https://github.com/user-attachments/assets/866d241c-8d63-4d49-9657-9b02737e9c40?width=100&height=100
If you find this code helpful in your research or wish to refer to the baseline results, please use the following BibTeX entry.
@article{shin2024enhancing,
title={Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting},
author={Shin, Inkyu and Yu, Qihang and Shen, Xiaohui and Kweon, In So and Yoon, Kuk-Jin and Chen, Liang-Chieh},
journal={arXiv preprint arXiv:2406.02541},
year={2024}
}
Python
96.2%
Cuda
2.2%
C++
1.1%
This repository contains the official Pytorch implementation of the paper "Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting".
According to our paper, we conducted two tasks with the following datasets.
There are two options for pre-processing the datasets.
We organize the datasets as follows:
├── datasets
│ | recon
│ ├── DAVIS
│ ├── JPEGImages
│ ├── 480p
│ ├── blackswan
│ ├── blackswan_pts_camera_from_deva
│ ├── ...
│ | edit
│ ├── DAVIS
│ ├── 480p_frames
│ ├── bmx-rider
│ ├── bmx-rider_pts_camera_from_deva
Setting up environments for training contains three parts:
git clone https://github.com/dlsrbgg33/Video-3DGS.git --recursive
cd Video-3DGS
conda create -n video_3dgs python=3.8
conda activate video_3dgs
# install pytorch
pip install torch==1.12.1+cu113 torchvision==0.13.1+cu113 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu113
# install packages & dependencies
bash requirement.sh
Setting up environments for evaluation contains two parts:
cd models/optical_flow/RAFT
bash download_models.sh
cd models/clipscore
git lfs install
git clone https://huggingface.co/openai/clip-vit-large-patch14
git clone https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K
bash sh_recon/davis.sh
To effectively obtain reprentation for video editing, we utilize all the training images for each video scene in this stage.
Arguments:
https://github.com/user-attachments/assets/8eb8e201-ef3b-461c-985b-72d3fa19cd54?width=100&height=100
bash sh_edit/{initial_editor}/{dataset}.sh
We currently support three "initial editors": Text2Video-Zero / TokenFlow / RAVE
We recommend user to install related packages and modules of above initial editors in Video-3DGS framework to conduct initial video editing.
For running TokenFlow efficiently (e.g., edit long video), we borrowed the some strategies from here.
https://github.com/user-attachments/assets/923dec4a-fb23-4c02-a187-a51cb57501a8?width=100&height=100
bash sh_edit/{initial_editor}/davis_re.sh
https://github.com/user-attachments/assets/866d241c-8d63-4d49-9657-9b02737e9c40?width=100&height=100
If you find this code helpful in your research or wish to refer to the baseline results, please use the following BibTeX entry.
@article{shin2024enhancing,
title={Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting},
author={Shin, Inkyu and Yu, Qihang and Shen, Xiaohui and Kweon, In So and Yoon, Kuk-Jin and Chen, Liang-Chieh},
journal={arXiv preprint arXiv:2406.02541},
year={2024}
}
Python
96.2%
Cuda
2.2%
C++
1.1%