This repository contains the necessary code to run the experiments in the paper "ExDDV: A New Dataset for Explainable Deepfake Detection in Video".
The dataset is published under the CC BY-NC-SA 4.0 license.
The csv file containing the annotations can be found in this repository. Use the columns movie_name and dataset to determine the input movie name.
There are three different folders inside the src folder containing the three models we used: LAVIS (BLIP-2), LLaVA and Phi3-Vision-Finetune. These are forked from the following repositories: LAVIS, LLaVA and Phi3-Vision-Finetune.
Each model is trained as per its corresponding repository. We used the respective training scripts: BLIP-2, LLaVA and Phi3-Vision. The only modification needed is to replaced the paths to the datasets in the corresponding files.
!Note for Phi3-Vision: The file processing_phi3_v.py from HuggingFace Transformers must be replaced with the script from here.
The current implementation uses LLaVA 1.5, BLIP-2 and PHI3-Vision.
Note:
The code for each model is provided in separate Jupyter notebooks:
llava_incontext.ipynbphi_incontext.ipynbblip_incontext.ipynbRequirements will differ based on the vision model used. We have followed the indications from the original paper. For LLAVA, for example, use the following steps:
git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
conda create -n llava python=3.10 -y
conda activate llava
pip install --upgrade pip # enable PEP 660 support
pip install -e .
For other vision models (e.g., BLIP-2, PHI3-Vision), refer to their official installation instructions.
Dataset CSV:
Ensure your dataset CSV (e.g., dataset_last.csv) includes columns such as movie_name, dataset, manipulation, movie_path, click_locations, and text. The CSV should be split into training, validation(not used here), and test sets.
Video Files:
Place your video files in the data/ folder. The directory structure should follow the dataset names (e.g., data/Faceforensics++/).
The pipeline outputs CSV files containing:
If you use this dataset or code in your research, please cite the corresponding paper:
Bibtex:
@article{hondru2025exddv,
title={ExDDV: A New Dataset for Explainable Deepfake Detection in Video},
author={Hondru, Vlad and Hogea, Eduard and Onchis, Darian and Ionescu, Radu Tudor},
journal={arXiv preprint arXiv:2503.14421},
year={2025}
}
Jupyter Notebook
83.1%
Python
16.6%
This repository contains the necessary code to run the experiments in the paper "ExDDV: A New Dataset for Explainable Deepfake Detection in Video".
The dataset is published under the CC BY-NC-SA 4.0 license.
The csv file containing the annotations can be found in this repository. Use the columns movie_name and dataset to determine the input movie name.
There are three different folders inside the src folder containing the three models we used: LAVIS (BLIP-2), LLaVA and Phi3-Vision-Finetune. These are forked from the following repositories: LAVIS, LLaVA and Phi3-Vision-Finetune.
Each model is trained as per its corresponding repository. We used the respective training scripts: BLIP-2, LLaVA and Phi3-Vision. The only modification needed is to replaced the paths to the datasets in the corresponding files.
!Note for Phi3-Vision: The file processing_phi3_v.py from HuggingFace Transformers must be replaced with the script from here.
The current implementation uses LLaVA 1.5, BLIP-2 and PHI3-Vision.
Note:
The code for each model is provided in separate Jupyter notebooks:
llava_incontext.ipynbphi_incontext.ipynbblip_incontext.ipynbRequirements will differ based on the vision model used. We have followed the indications from the original paper. For LLAVA, for example, use the following steps:
git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
conda create -n llava python=3.10 -y
conda activate llava
pip install --upgrade pip # enable PEP 660 support
pip install -e .
For other vision models (e.g., BLIP-2, PHI3-Vision), refer to their official installation instructions.
Dataset CSV:
Ensure your dataset CSV (e.g., dataset_last.csv) includes columns such as movie_name, dataset, manipulation, movie_path, click_locations, and text. The CSV should be split into training, validation(not used here), and test sets.
Video Files:
Place your video files in the data/ folder. The directory structure should follow the dataset names (e.g., data/Faceforensics++/).
The pipeline outputs CSV files containing:
If you use this dataset or code in your research, please cite the corresponding paper:
Bibtex:
@article{hondru2025exddv,
title={ExDDV: A New Dataset for Explainable Deepfake Detection in Video},
author={Hondru, Vlad and Hogea, Eduard and Onchis, Darian and Ionescu, Radu Tudor},
journal={arXiv preprint arXiv:2503.14421},
year={2025}
}
Jupyter Notebook
83.1%
Python
16.6%