Parul-Gupta/MultiFakeVerse

A pipeline for automatic and context-relevant image tampering. (ACM Multimedia 2025)

5

stars

18

commits

Python

primary language

Aug 12, 2025

updated

README

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations

A pipeline for automatic and semantic context-relevant image tampering.

Dataset available on Huggingface Preview Version: Link

https://github.com/user-attachments/assets/5743d11d-a201-443c-b546-efc8b6e5157e

Abstract

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-scale and reasoning capability driven deepfake benchmark dataset specifically tailored for person-centric object, context and scene manipulations. In this paper, we address this gap by introducing MultiFakeVerse, a large scale person-centric deepfake dataset, comprising 845,286 images generated through manipulation suggestions and image manipulations both derived from vision-language models (VLM). The VLM instructions were specifically targeted towards modifications to individuals or contextual elements of a scene that influence human perception of importance, intent, or narrative. This VLM-driven approach enables semantic, context-aware alterations such as modifying actions, scenes, and human-object interactions rather than synthetic or low-level identity swaps and region-specific edits that are common in existing datasets. Our experiments reveal that current state-of-the-art deepfake detection models and human observers struggle to detect these subtle yet meaningful manipulations.

Image MultiFakeVerse. A brief overview of the proposed dataset. Here, we introduce subtle and profound person-centric deepfakes covering person-level, object-level, scene-level, (person+object)-level, (person+scene)-level manipulations. Image best viewed in color.

Some more examples

Image

Image Analyzing the perceptual impact of manipulations in images. The edited regions are highlighted by yellow boxes. The analysis covers changes in attributes such as perceived emotion, identity, and ethical implications.

Image Some example images and their fakes obtained using VLM based image editing, The highlighted yellow boxes indicate the edited regions for images with localized edits. The rest of the images have all three kinds of modifications: Person-level, Object-level and Scene-level.

Contributors

Parul-Gupta

18 commits

Parul-Gupta/MultiFakeVerse

A pipeline for automatic and context-relevant image tampering. (ACM Multimedia 2025)

5

stars

18

commits

Python

primary language

Aug 12, 2025

updated

README

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations

A pipeline for automatic and semantic context-relevant image tampering.

Dataset available on Huggingface Preview Version: Link

https://github.com/user-attachments/assets/5743d11d-a201-443c-b546-efc8b6e5157e

Abstract

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-scale and reasoning capability driven deepfake benchmark dataset specifically tailored for person-centric object, context and scene manipulations. In this paper, we address this gap by introducing MultiFakeVerse, a large scale person-centric deepfake dataset, comprising 845,286 images generated through manipulation suggestions and image manipulations both derived from vision-language models (VLM). The VLM instructions were specifically targeted towards modifications to individuals or contextual elements of a scene that influence human perception of importance, intent, or narrative. This VLM-driven approach enables semantic, context-aware alterations such as modifying actions, scenes, and human-object interactions rather than synthetic or low-level identity swaps and region-specific edits that are common in existing datasets. Our experiments reveal that current state-of-the-art deepfake detection models and human observers struggle to detect these subtle yet meaningful manipulations.

Image MultiFakeVerse. A brief overview of the proposed dataset. Here, we introduce subtle and profound person-centric deepfakes covering person-level, object-level, scene-level, (person+object)-level, (person+scene)-level manipulations. Image best viewed in color.

Some more examples

Image

Image Analyzing the perceptual impact of manipulations in images. The edited regions are highlighted by yellow boxes. The analysis covers changes in attributes such as perceived emotion, identity, and ethical implications.

Image Some example images and their fakes obtained using VLM based image editing, The highlighted yellow boxes indicate the edited regions for images with localized edits. The rest of the images have all three kinds of modifications: Person-level, Object-level and Scene-level.

Contributors

Parul-Gupta

18 commits

Languages

Python

100.0%