ant-research/Awesome-AIGC-Image-Video-Detection

A curated collection of the latest research and resources on AI-Generated Image and Video Detection.

See the code

README

Awesome AIGC Image/Video Detection Awesome

Overview

A curated collection of the latest research and resources on AI-Generated Image and Video Detection. This repository encompasses datasets, benchmarks, research papers, and practical detection tools.

🚀🚀🚀Contributions are welcome! If you find any missing papers, datasets, or tools, feel free to open an issue or submit a pull request.

Contents


🔥 Hot Events


Benchmarks & Datasets

Modality Legend: [I] Image | [V] Video | [M] Multi-modal

Annotation Type Legend: Au: Authenticity | Ex: Explainability | Lo: Localization

BenchmarkPaperVenue & YearModalityNotesReal SourceFake Source/GeneratorAnnotationScaleDownload
DF26DF26: We Cannot Tell Fake From Real AnymoreArxiv 2026[V]Modern-Generator AIGV Benchmark, Single-Person Public-Speaking Scenarios (Direct-to-Camera / Official Statement / Studio Interview), per-clip scene-prompt control (identity & scene fixed, generator varied), human + SOTA detector study near random chance, closed-source subset evaluation-onlyOpenVid-1M, TalkingCelebs, MAVOS-DD7 recent T2V/I2V models: Wan2.6, Veo 3.1, Grok Imagine 1.0, Kling 3.0 (closed); Wan2.2-A14B, HunyuanVideo 1.5, LTX 2.3 (open)Au2.7K (271 real + 2,420 fake, 1280×720)DF26 (controlled access)
DailyBenchDailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative ModelsArxiv 2026 (v3 2026-09-01)[I]Unified AIID testbed = FakeBench (full T2I synthesis) + ManipulationBench (object-level edits on real images); LAION-Aesthetics V2 pool filtered by aesthetic score ≥6 & shortest side ≥512; per-generator real/fake pairs; robustness study under recompression & pixel perturbation; ships FPD diagnostic baselineLAION-Aesthetics V2T2I: SD3.5-Large, FLUX.1, FLUX.2, Qwen-Image-2512, Z-Image, Nano Banana 2, GPT-Image 2; Edit: FLUX-Fill (random/object mask), FLUX.2-klein-9B, Qwen-Image-Edit-2511, Step1X-Edit-v1p2Au270K(v3:FakeBench ≈185K + ManipulationBench ≈75K)DailyBench
Project
AGIDefect-4KAGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and ExplanationACM MM 2026[I]Defect Detection, Localization & Explanation, 15 SOTA Generators, Quality ScoringDALL-E 3, Midjourney, FLUX, Gemini, GPT-Image, Ideogram, Kling, Grok, etc.Au, Lo, Ex4KAGIDefect-4K
RA-BenchCan We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social DisseminationArxiv 2026[V]Real-Crisis-Anchored, Source-Matched Evaluation, Human-Proof Subset, Propagation RobustnessReal crisis event footage (public media & source URLs)4 open-source + 5 closed-source generators (incl. Wan2.2)Au17.9KRA-Bench
RealHDRealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated ImagesArxiv 2026[I]Multi-category, SOTA Generators, 10K+ Prompts, Inpainting MasksT2I, Inpainting, Refinement, Face SwappingAu, Lo730K+RealHD
TreasureFleet: Few Shots Lead Effective AI-generated Image DetectionICML 2026[I]64 Models, 20 Closed-source Commercial Engines, Few-shot AdaptationDiverse architectures & 20 commercial enginesAu360KTreasure
LADBenchLADBench: A Benchmark for Logical Fault Detection in ImagesICDL 2026[I]Logical Anomaly Detection, VLM Evaluation, Common Sense ReasoningSynthetic images with logical anomaliesAu1K+LADBench
EVID-BenchWhen Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation DetectionArxiv 2026[V]Search-Grounded Verification, Evidence-Dependent ManipulationAuEVID-Bench
CoCoVideoCoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video DetectionArxiv 2026[V]Commercial AIGC Models, Contrastive BenchmarkCommercial video generation modelsAuCoCoVideo
FraudBenchFraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund EvidenceArxiv 2026[M]Fraudulent Refund Detection (ECommerce)AuFraudBench
GPT-Image-2 WildGPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of DeploymentArxiv 2026[I]GPT-Image-2Twitter (real images)GPT-Image-2Au10KGPT-Image-2 Wild
Artifact-BenchArtifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated VideosArxiv 2026[V]MLLM Evaluation, Video ArtifactsAuArtifact-Bench
CommGen15PGC: Peak-Guided Calibration for Generalizable AI-Generated Image DetectionICML 2026[I]15 Commercial Generative Models15 commercial generatorsAuCommGen15
AEGIS-AcademicAEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic ImagesArxiv 2026[I]Academic Image ForensicsAuAEGIS
SciFigDetectSciFigDetect: A Benchmark for AI-Generated Scientific Figure DetectionArxiv 2026[I]Scientific Figure DetectionNano Banana Pro, GPT-image-1.5Au150KSciFigDetect
ActivityForensicsActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosCVPR 2026[V]Action-level AIGC in videosAu6KActivityForensics
MintVidVideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningICML 2026[V]OpenVid, VFHQ, HDTF, TikTokJimeng3.0-Pro, Seedance, Kling2.5-Turbo, Sora2, TikTok, Youtube, etc.Au4KMintVid
AIGVDBenchYour One-Stop Solution for AI-Generated Video DetectionCVPR 2026[V]OpenVid-HD31 generation modelsAu440kAIGVDBench
HydraFakeVeritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningICLR 2026(Oral)[I]FFHQ, VFHQ, CelebAHQ, FF++, etc.GPT-4o, HailuoAI, ICLight, InfiniteYou, etc.Au, Ex100KHydraFake
BR-GenZooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification ApproachAAAI 2026[I]Au, Lo150KBR-Gen
RRDatasetBridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging ScenariosICCV 2025[I]Real-World Robustness, Internet Transmission, Re-digitizationAuDownLoad
HiResolutionNo Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image DetectionICLR 2026[I]Au50KHiRes-50K
AIGI-NowAlignGemini: Generalizable AI-Generated Image Detection Through Task-Model AlignmentArxiv 2026[I]COCONano Banana, GPT-4o, Jimeng, Kling, Minimax, etc.Au18KAIGI-Now
RealChainBeyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image DetectionArxiv 2026[I]Au14KRealChain
GenVidBenchGenVidBench: A 6-Million Benchmark for AI-Generated Video DetectionAAAI 2026[V]Au6MGenVidBench
SkyraSkyra: AI-Generated Video Detection via Grounded Artifact ReasoningCVPR 2026[V]Au, Ex, Lo4KViF-CoT-4K
So-Fake-SetSo-Fake: Benchmarking and Explaining Social Media Image Forgery DetectionArxiv 2025[I]F30k, WIDER, FFHQ, CelebA, OpenImages, COCO, OpenForensicsQwen-image, GPT-4o, Nano Banana, Seedream3.0, Ideogram3.0, etc.Au2M+So-Fake-Set
So-Fake-OOD
GenBuster++BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLMArxiv 2025[M]Au4KGenBuster++
GenBusterBusterX: MLLM-Powered AI-Generated Video Forgery Detection and ExplanationArxiv 2025[I]Au200KGenBuster-200K
AIGIBenchIs Artificial Intelligence Generated Image Detection a Solved Problem?NeurIPS 2025[I]FFHQ, CelebA-HQ, Open Images V7Common generators & SocialRF, CommunityAIAu200KAIGIBench
Ivy-FakeIVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC DetectionArxiv 2025[M]Au, Ex150KIvy-Fake
AEGISAEGIS: Authenticity Evaluation Benchmark for AI-Generated Video SequencesACM MM 2025[V]Vript (YouTube, TikTok), DVF, YouTube (self-collected)Stable Video Diffusion, CogVideoX-5B, I2VGen-XL, Pika, KLing, SoraAu, Ex10K+AEGIS
NeXT-IMDLNeXT-IMDL: Build Benchmark for Next-Generation Image Manipulation Detection & LocalizationArxiv 2025[I]Flickr30k, COCO, OpenImages V7SD2-Inpainting, SDXL-Inpainting, FLUX-Inpainting, etc.Au, Lo558KNeXT-IMDL
ARForensicsD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionICCV 2025[I]ImageNetInfinity, Janus_Pro, RAR, Switti, VAR, LlamaGen, Open_MAGVIT2Au300kARForensics
OpenSDIOpenSDI: Spotting Diffusion-Generated Images in the Open WorldCVPR 2025[I]Megalith-10MSD1.5, SD2.1, SDXL, SD3, Flux.1Au, Lo300KOpenSDI
Community ForensicsCommunity Forensics: Using Thousands of Generators to Train Fake Image DetectorsCVPR 2025[I]LAION, ImageNet, COCO, FFHQ, CelebA, MetFaces, AFHQ, etc.4803 generators (Latent Diffusion, GAN, Autoregressive, Pixel Diffusion, Commercial)Au2.7MCommunity Forensics
FakeClueSpot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationNeurIPS 2025[I]Au, Ex100KFakeClue
XAIGID-RewardBenchExplainable AI-Generated Image Detection RewardBenchNeurIPS 2025 Workshop[I]COCO-2017Imagen 4, Flux.1 Dev, Bagel, etc.Au, Ex3KXAIGID-RewardBench
RewardDataLearning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMsArxiv 2025[V]Au, Ex4.3KRewardData
OpenFakeOPENFAKE: An Open Dataset and Platform Toward Real-World Deepfake DetectionArxiv 2025[I]LAION-400MSD 1.5/2.1/XL/3.5, Flux 1.0-dev/1.1-Pro/Schnell, Midjourney v6/v7, DALL·E 3, Imagen 3/4, GPT Image 1, Ideogram 3.0, Grok-2, HiDream-I1, Recraft v3, Chroma, and 10 community LoRA/finetune variantsAu~4MOPENFAKE
Video Reality TestVideo Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?Arxiv 2025[V]YouTube ASMR (social media)Veo3.1-Fast, Sora2, Wan2.2-A14B, Wan2.2-5B, OpenSora-V2, HunyuanVideo, StepVideoAu149 real + dynamic fakeVideo Reality Test
DDLDDL: A Dataset for Interpretable Deepfake Detection and Localization in Real-World ScenariosArxiv 2025[M]Au367KDDL
DiffSeg30kDiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC DetectionArxiv 2025[I]COCOSD2, SD3.5, SDXL, Flux.1, Glide, Kolors, HunyuanDiT1.1, Kandinsky 2.2Au, Lo30KDiffSeg30k
FakePartsFakeParts: a New Family of AI-Generated DeepFakesArxiv 2025[V]Au, Lo81KFakeParts
ForensicHubForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and LocalizationNeurIPS 2025[I]ProGAN, StyleGAN, LDM, SDv1.4, SDv1.5, SDv2, SDXL, SD-ControlNet, MidJourney, ADM, GLIDE, VQDM, BigGANAu, Lo23 datasets
42 models
ForensicHub
LOKILOKI: A Comprehensive Synthetic Data Detection Benchmark Using Large Multimodal ModelsICLR 2025[M]SORA, Keling, Open-Sora, FLUX, Midjourney, Stable Diffusion, Nerf-based, Gaussian-based, GPT-4o, Qwen-Max, Llama 3.1-405B, MusicGen, AudioLDM2...Au, Ex18KLOKI
ChameleonA Sanity Check for AI-Generated Image DetectionICLR 2025[I]UnsplashMidjourney, DALLE-3, Stable Diffusion (various LoRA fine-tuned)Au26KChameleon
WildFakeWildFake: A Large-scale Challenging Dataset for AI-Generated Images DetectionAAAI 2025[I]Au3.7MWildFake
WildRFReal-Time Deepfake Detection in the Real-WorldArxiv 2024[I]Reddit, X (Twitter), Facebook (real images)Reddit, X (Twitter), Facebook (social media deepfakes)AuWidlRF
AIGCDetectBenchmarkPatchCraft: Exploring Texture Patch for Efficient AI-generated Image DetectionArxiv 2024[I]Au100KAIGCDetectionBenchMark
GenVideoDeMamba: AI-Generated Video Detection on Million-Scale GenVideo BenchmarkArxiv 2024[V]Au2.3MGenVideo
DRCTDrct: Diffusion reconstruction contrastive training towards universal detection of diffusion generated imagesICML 2024[I]MSCOCOLDM, SDv1.4, SDv1.5, SDv2, SDXL, SD-ControlNetAu2MDRCT-2M
GenImageGenImage: A Million-Scale Benchmark for Detecting AI-Generated ImageNeurIPS 2023[I]ImageNet, WukongMidJourney, SDv1.4, SDv1.5, ADM, GLIDE, VQDM, BigGANAu2.7MGenImage
DF40DF40: Toward Next-Generation Deepfake DetectionNeurIPS 2024[I] [V]Au0.1M+ videos, 1M+ imagesDF40
Forensics-BenchForensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language ModelsCVPR 2025[I], [V], [M]Various public datasetsGAN, Diffusion, VAE, RNN, Encoder-Decoder, Graphics-basedAu, Lo63KForensics-Bench

⬆ Back to Top


Research Papers

💡 Note: Papers are sorted by year (descending) within each category.
Modality Legend: [I] Image | [V] Video | [M] Multi-modal
Category Layout: Top level is split into MLLM-Based (MLLM-powered detection) and Classification-Based (compact/traditional classifiers). Classification-Based is further organized into six subcategories. When a paper fits multiple subcategories, the priority is: Training-Free/Zero-Shot > Continual/Incremental > Video Spatiotemporal > Frequency/Low-Level Artifacts > Supervised General.

MLLM-Based

This category focuses on utilizing Multimodal Large Language Models (MLLMs) like GPT-4V, LLaVA, or Qwen-VL to detect AI-generated content. These methods often provide natural language explanations (explainability) alongside binary detection.

TitleVenue & YearModalityHighlights/KeywordsCode
Evidence-Guided Detection, Localization and Explanation for Text-Centric Image ForensicsACM MM 2026 Challenge[I]Detector-Localizer-Reasoner Cascade, Iterative Difficulty-Aware Mining, Report-Mask ConsistencyGitHub
AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and ExplanationACM MM 2026[I]AGIDefect-4K Dataset, Hierarchical Defect Annotation, MLLM Baseline (AGIDA)GitHub
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation OptimizationACM MM 2026[I]Feature-robust Augmentation, Mean-Teacher Consistency, Evidence-grounded Preference Optimization, Challenge WinnerGitHub
PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMsIJCAI 2026 Workshop[I]Perception-as-Tool, DINOv3 Forensic Perception Tool, General-Purpose MLLM ExplanationGitHub
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI DetectionACM MM 2026[I]Interactive Visual Search, Verifier-guided Evidence Alignment, GroundFake Dataset, FakeFrontier BenchmarkN/A
SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited DataArxiv 2026[I]Adversarial RL Loop, Diffusion Editor, Free-form ExplanationN/A
VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsArxiv 2026[V]Meta-Detection RL, Verifiable Temporal Grounding, Evidence-Guided Reward RedistributionN/A
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI DetectionArxiv 2026[I]Value-aware On-Policy Distillation, Perception-Enhanced ReasoningGitHub
Detecting AI-Generated Video: A Vision-Language Dual-View SurveyACL 2026 Findings[V][Survey] Vision-Language Dual-View Taxonomy, Factual Fidelity Verification, Cross-modal ConsistencyN/A
TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image DetectionICML 2026[I]Artifact feature and Semantic feature fusionGitHub
Venus-DeFakerOne: Unified Fake Image Detection & LocalizationArxiv 2026[I]Unified Detection & Localization, Large-Scale TrainingGitHub
GenShield: Unified Detection and Artifact Correction for AI-Generated ImagesICML 2026[I]Unified MLLM, Detect & Correct ArtifactsGitHub
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationArxiv 2026[I]Reasoning-Aligned Representation, InterpretableN/A
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image DetectionCVPR 2026[I]Unified Model, Co-Evolution (Generation & Detection)GitHub
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningICML 2026[V]Perception Pretext RL, Fact-based Reasoning, MintVid DatasetGitHub
Veritas: Generalizable deepfake detection via pattern-aware reasoningICLR 2026(Oral)[I]Pattern-aware Reasoning, HydraFake DatasetGithub
DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-ReflectionArxiv 2026[I]Knowledge Injection, Self-ReflectionN/A
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic ReasoningArxiv 2026[M]Agentic Framework, Document SafetyN/A
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RLICLR 2026[V]Multi-stage RL (GRPO), Time Artifacts, Video Detection DatasetGitHub
FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningICLR 2026[I]Grounded Reasoning, Human-annotated DatasetN/A
AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model AlignmentArxiv 2026[I]Decoupling (Semantic & Pixel), AIGI-Now DatasetN/A
Zoom-In to Sort AI-Generated Images OutArxiv 2026[I]Thinking with Images, MagniFake DatasetN/A
AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image DetectionArxiv 2026[I]Agentic frameworkGithub
EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image DetectionArxiv 2026[I]Agentic Framework, Method EnsemblingN/A
VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake DetectionArxiv 2026[I]Part-centric Forensic, OmniFake DatasetProject
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?Arxiv 2026[V]GenVideoLens benchmarkN/A
Semantic Visual Anomaly Detection and Reasoning in AI-Generated ImagesICLR 2026[I]Semantic Anomaly Reasoning, AnomReason DatasetN/A
FAKE-HR1: RETHINKING REASONING OF VISION LANGUAGE MODEL FOR SYNTHETIC IMAGE DETECTIONArxiv 2026[I]Hybrid-Reasoning, Dual-mode DatasetN/A
MIRAGE: Towards AI-Generated Image Detection in the WildArxiv 2025[I]Human Curation Dataset, Heuristic-to-Analytic ReasoningN/A
BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLMArxiv 2025[M]RL Post-training, Cross-Modal, Thinking Reward MechanismGithub
BusterX: MLLM-Powered AI-Generated Video Forgery Detection and ExplanationArxiv 2025[V]GenBuster-200K Dataset, Cold Start + RL TrainingGithub
REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image DetectionArxiv 2025[I]Chain-of-Evidence, Expert-grounded RLN/A
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationNeurIPS 2025[I]FakeClue Dataset, Fine-grained Artifact Clues, Artifact ExplanationGitHub
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsICCV 2025[I]Holmes-Set, Multi-Expert Jury, 3-Stage Training PipelineGithub
LEGION: Learning to Ground and Explain for Synthetic Image DetectionICCV 2025[I]SynthScars Dataset, Defender & Controller, Image RefinementGitHub
Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image DetectionArxiv 2025[I]Perception & Reasoning, ExplainFake-BenchN/A
SIDA: Social Media Image Deepfake Detection, Localization, and ExplanationCVPR 2025[I]SID-Set, Mask Prediction, Social Media ContextGithub
FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language ModelsICLR 2025[I]Explainable IFDL, Domain Tag-guided, Multi-modal LocalizationGitHub
FakeScope: Large Multimodal Expert Model for Transparent AI-Generated Image ForensicsArxiv 2025[I]FakeChain Dataset, FakeInstruct, Trace EvidenceN/A
AntifakePrompt: Prompt-Tuned Vision-Language Models are Fake Image DetectorsArxiv 2024[I]VQA, InstructBLIP, Soft Prompt-tuning, Zero-shotGitHub

Classification-Based

This category includes supervised learning approaches that train neural networks (CNNs, ViTs, VFMs, etc.) specifically to classify authentic vs. AI-generated content. They usually focus on robustness, generalization, and feature extraction. It is organized into six subcategories: Supervised General Detectors, Video Spatiotemporal Modeling, Training-Free / Zero-Shot, Frequency-Domain & Low-Level Artifacts, Continual & Incremental Learning, and Related & Other.

Supervised General Detectors

Trainable classifiers and backbones (CNNs, ViTs, vision foundation models, CLIP-based adapters, few-shot and prompt-based methods) for general-purpose detection.

TitleVenue & YearModalityHighlights/KeywordsCode
Learning Continuous Source Responses For Generalizable AI-Generated Image DetectionArxiv 2026[I]CuRe, Continuous Source-Response Regression (Real-Generated Mixing Ratio), Shortcut-Cue Suppression, Source-Response Subspace, 10-Benchmark EvaluationGitHub
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image DetectionArxiv 2026[I]UCF-Net, CLIP Semantic + DINO Structural Priors, Layer-wise Expert Aggregation, Entropy-based Uncertainty Fusion, 4M-image Unified Benchmark, Cross-generator EvaluationGitHub
FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and LocalizationArxiv 2026[I]Forensic-Semantic MoE, Joint Detection & Localization, Cross-generator Generalization (OpenSDID)GitHub
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation LocalizationArxiv 2026[I]Global Artifact Token, SAM3 FiLM Injection, Boundary Adhesion AnalysisN/A
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic ResidualsECCV 2026 (Spotlight)[I]Low-Rank Collapse Signature, Semantic-Residual Decoupling, Cross-model GeneralizationN/A
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image DetectionECCV 2026[I]Closed-form Gaussian Heads, Percept-Lens Protocol (39 Datasets), Transfer DiagnosticN/A
Environment-Invariant Subspace Learning for Generalizable Deepfake DetectionArxiv 2026[I]Environment-Invariant Subspace, VFM Semantic Priors, Environmental InterventionN/A
Understanding Why Foundation Models Work for Diffusion-Generated Image DetectionArxiv 2026[I]Interpretability Analysis, DDIM Inversion, Low-to-Mid Frequency Distributional DiscrepancyN/A
PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image DetectionArxiv 2026[I]DINO Patch Tokens, 2D Spatial Aggregation, LoRA AdaptersN/A
GlobalForge: Towards Robust AI-Generated Image DetectionArxiv 2026[I]Global Structural Reasoning, Local Information Bottleneck, RealDeg-BenchCode
Fleet: Few Shots Lead Effective AI-generated Image DetectionICML 2026[I]Few-shot Adaptation, Routing Correction, Treasure BenchmarkGitHub
SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision EncodersArxiv 2026[I]Frozen Vision Encoders, Linear Classifier, RealWorldBenchN/A
HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image DetectionACM MM 2026[I]Asymmetric Prompting, Dynamic Decision BoundaryN/A
VINA: Video as Natural Augmentation: Towards Unified AI-Generated Image and Video DetectionArxiv 2026[M]Unified Image/Video Detection, Cross-Modal Contrastive LearningN/A
PGC: Peak-Guided Calibration for Generalizable AI-Generated Image DetectionICML 2026[I]Peak-Guided Calibration, CommGen15 DatasetGitHub
Reduce the Artifacts Bias for More Generalizable AI-Generated Image DetectionArxiv 2026[I]Bias-free TrainingGitHub
Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification ApproachAAAI 2026[I]Localized AIGC Detection, Forgery Amplification, Scene-aware Local ForgeryGitHub
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation ModelsArxiv 2026[I]Linear Probe, Vision Foundation Models, Emergent Forensic CapabilityN/A
MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image DetectionArxiv 2026[I]Manifold Reconstruction, Memory Bank, Human-AIGI BenchmarkGitHub
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image DetectionICLR 2026[I]Detail-preserving dual-path architecture, Multi-task learning, HiRes-50K benchmarkN/A
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch LearningICLR 2026[I]Random Patch Replacement, Patch-wise Contrastive LearningN/A
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the WildArxiv 2025[I]Mixture-of-Experts, Semantic-Artifact Decoupling, Mirage DatasetGitHub
DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image DetectionArxiv 2025[I]Blur Robustness, Knowledge Distillation, DINOv3Github
Orthogonal Subspace Decomposition for Generalizable AI-Generated Image DetectionICML 2025 (Oral)[I]SVD Orthogonal Subspace, Asymmetry Phenomenon, Parameter-efficient Fine-tuningGitHub
A Bias-Free Training Paradigm for More General AI-generated Image DetectionCVPR 2025[I]Bias-Free, Semantic Alignment, Stable Diffusion Self-conditioningGithub
Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionCVPR 2025[I]CLIP, Blending Boundaries, Forgery-aware Prompt LearningGithub
Exploring Unbiased Deepfake Detection via Token-Level Shuffling and MixingAAAI 2025[I]Token-Level Shuffling, Contrastive Loss, Bias MitigationN/A
FakeFormer: Efficient Vulnerability-Driven Transformers for Generalisable Deepfake DetectionArxiv 2024[I]Vulnerability-driven, Local Attention (L2-Att), Vision TransformerGitHub

Video Spatiotemporal Modeling

Methods that exploit temporal inconsistencies, motion patterns, and spatiotemporal artifacts in AI-generated videos.

TitleVenue & YearModalityHighlights/KeywordsCode
Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video DetectionACM MM 2026[V]Cross-Scale Coupling Mismatch, Macro Temporal Dynamics vs. Pixel-Level Residuals, Persistent Homology, Encoder-AgnosticGitHub
MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow TrajectoriesArxiv 2026[V]Physical Motion Consistency, Sparse Optical-Flow Trajectories, Multi-scale Geometric EvolutionN/A
Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video DetectionArxiv 2026[V]V-PVP Readout, Patch Velocity Profiling, Frozen Video BackbonesCode
Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video DetectionArxiv 2026[V]Motion Bias Analysis, Preprocessing/Sampling Bias, Frequency-based ComparisonN/A
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementArxiv 2026[V]Counterfactual Intervention, Causal Disentanglement, Cross-domain GeneralizationGitHub
ReConFuse: Reconstruction-Error Guided Semantic Fusion for AI-Generated Video DetectionArxiv 2026[V]Reconstruction Error, Semantic Fusion, Spatial-Temporal ArtifactsN/A
Detecting AI-Generated Videos with Spiking Neural NetworksArxiv 2026[V]Spiking Neural Networks, Temporal ArtifactN/A
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video DetectionArxiv 2026[V]Cross-Modal Temporal Artifacts, Video DetectionN/A
Preserving Forgery Artifacts: AI-Generated Video Detection at Native ScaleICLR 2026[V]Native scale video processing, Massive realistic video dataset, Preserves subtle generation artifactsN/A
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationNeurIPS 2025[V]Wavelet-band Augmentation, Forensic Frequency Artifacts, Single-generator GeneralizationGitHub
AI-Generated Video Detection via Perceptual StraighteningNeurIPS 2025[V]Perceptual Straightening, DINOv2, Temporal CurvatureGitHub
Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionNeurIPS 2025[V]Normalized Spatiotemporal Gradient (NSG), Maximum Mean Discrepancy (MMD)Github
Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated ContentCVPR 2025[V]SigLIP-So400M, Attention-Diversity Loss, Full-frame ManipulationsN/A
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake DetectionTMM 2025[V]Direction-aware Attention, SpatioTemporal Invariant LossN/A
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo BenchmarkArxiv 2024[V]Mamba, State Space Model, Long-range Spatiotemporal InconsistencyGitHub
Distinguish Any Fake Videos: Unleashing the Power of Large-scale Data and Motion FeaturesArxiv 2024[V]GenVidDet, Optical Flow, Dual-Branch 3D TransformerN/A

Training-Free / Zero-Shot

Methods that detect AI-generated content without additional training on detection data.

TitleVenue & YearModalityHighlights/KeywordsCode
Frozen DINO Localizes Image Edits Without a LocalizerArxiv 2026[I]Training-free, Frozen DINO Patch-token Drift, Haar Perturbation, Edit LocalizationGitHub
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal RoughnessECCV 2026[V]Training-free, Patch-level Incoherence, Temporal Roughness, Ultra-low FPRGitHub
Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsCVPR 2026[V]Training-free, Zero-shot, Spatial-Temporal Likelihoods, ComGenVid DatasetGitHub

Frequency-Domain & Low-Level Artifacts

Methods based on spectral analysis, quantization/upsampling traces, and other low-level generative artifacts.

TitleVenue & YearModalityHighlights/KeywordsCode
Structured Local Differential Modeling for AI-Generated Image DetectionArxiv 2026[I]RippleNet, Local Differential Signals, Low-SNR Forgery TracesN/A
Dual Data Alignment Makes AI-Generated Image Detector Easier GeneralizableNeurIPS 2025 (Spotlight)[I]Dual-domain Alignment, Frequency-level Bias, VAE ReconstructionGitHub
D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionICCV 2025[I]Discrete Distribution Discrepancy-aware Transformer, Vector Quantized Variational AutoEncoderGithub
Any-Resolution AI-Generated Image Detection by Spectral LearningCVPR 2025[I]Spectral Context Attention, Frequency Reconstruction, OOD DetectionGithub
Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningAAAI 2024[I]Frequency Domain, FFT, Frequency Conv Layer (FCL), LightweightGitHub
Rethinking the Up-Sampling Operations in CNN-based Generative NetworkCVPR 2024[I]Neighboring Pixel Relationships, Generalized Structural ArtifactsGithub

Continual & Incremental Learning

Methods that keep adapting detectors to evolving generators without catastrophic forgetting.

TitleVenue & YearModalityHighlights/KeywordsCode
Preserving Knowledge across Space and Time for Continual Video Deepfake DetectionECCV 2026[V]Modality-Specific Frequency Distillation, Spatial/Temporal/Spatiotemporal Decomposition, Cross-Modality DecorrelationGitHub
Automated In-the-Wild Data Collection for Continual AI Generated Image DetectionArxiv 2026[I]Continual Learning, Continual Data CollectionGitHub
IncreFA: Breaking the Static Wall of Generative Model AttributionArxiv 2026[I]Incremental Learning, Generative Model AttributionGitHub
SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual LearningCVPR 2026[I]Scene-aware optimization, Continual learningGitHub
Generalizable and Adaptive Continual Learning Framework for AI-generated Image DetectionTMM 2026[I]Continual Learning, Kronecker-Factored Approximate CurvatureN/A

Papers related to AI-generated content safety (e.g., provenance/watermarking, misinformation verification) that do not fit the subcategories above.

TitleVenue & YearModalityHighlights/KeywordsCode
DF26: We Cannot Tell Fake From Real AnymoreArxiv 2026[V][Benchmark] 2,691 Public-Speaking Videos, 7 Modern T2V/I2V Models, Human & SOTA Detectors Near Chance, Distribution-Shift RobustnessN/A
APT: Anchor-aligned Perturbations for Tamper Localization in Fully Regenerated ImagesECCV 2026[I][Proactive Forensics] Semi-Fragile Latent Perturbation, Fully Regenerated (Inpainting) Setting, Anchor-Direction AlignmentN/A
Training-Free Reconstruction-Based AI-Generated Image Detectors Are Inherently Vulnerable to Adversarial ExamplesECCV 2026 Workshop[I][Robustness Analysis] Reconstruction-based Detector Attacks, Transferable Adversarial Examples, Real-world DegradationsN/A
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social DisseminationArxiv 2026[V][Evaluation] RA-Bench, Crisis Event Videos, Detector Generalization, Social DisseminationN/A
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation DetectionArxiv 2026[V]Search-Grounded Verification, EVID-Bench, Evidence-Dependent ManipulationN/A
Robust ASIC-Based Image Authentication Using Reed-Solomon LSB WatermarkingPreprint 2026[I]ASIC PoW, Hardware-bound Provenance, Reed-Solomon WatermarkingGitHub

⬆ Back to Top


Competitions

CompetitionLinkYearInfo
Robust AIGC DetectionNTIRE 2026 Robust AI-Generated Image Detection in the Wild2026No restrictions on training data.
Evaluate ROC AUC metrics on robust samples.
Robust Deepfake DetectionNTIRE 2026 Robust Deepfake Detection Challenge2026No restrictions on training data.
The 6th Face Anti-spoofing ChallengeThe 6th Face Anti-Spoofing: Unified Physical-Digital Attacks Detection@ICCV20252025No external data or pre-trained models allowed.
Limited to a single DL model with under 100G FLOPs.
Detect AI vs. Human-Generated Images2025 Women in AI (WAI) Kaggle Challenge2025Paired dataset of authentic and AI-generated images
The 5th Face Anti-spoofing Challenge5th Chalearn Face Anti-spoofing Workshop and Challenge@CVPR20242024UniAttackData+ for unified physical and digital attack detection.

⬆ Back to Top


Practical Detection Tools

  • 美亚鉴真 - 微信小程序搜索 美亚鉴真
  • SiliconSignature - GitHub - Hardware-bound image authentication using ASIC PoW nonces for unforgeable provenance certification
  • EyeSift - Website - Free online AI text/image/video/audio detector with detailed per-model benchmarks
  • Hive Moderation - Website
  • Tencent Zhuque AI Detection Assistant - Website
  • AI or Not - Website
  • Illuminarty - Website
  • Winston AI - Website
  • Is it AI? - Website
  • TruthScan - Website
  • 中科睿鉴 (Zhongke Ruijian) - 微信小程序搜索 睿鉴AI

⬆ Back to Top


🏢 About Our Team

We are the Content Security Intelligence Team under Ant Group - Machine Intelligence. We are responsible for developing comprehensive content security and risk-mitigation capabilities for the Ant Group ecosystem, bridging the gap between rapidly evolving technologies and the urgent need for digital trust.

Why We Do It

In an era where synthetic media is increasingly sophisticated and pervasive, our research serves as a critical line of defense. By advancing AIGC detection technologies, we aim to:

  • Safeguard Digital Integrity: We provide essential defense mechanisms to protect the authenticity of visual content and combat the spread of misinformation in the digital space.
  • Empower Trust: Our solutions ensure the public can distinguish between genuine and synthetic media, fostering a more transparent and trustworthy digital ecosystem.
  • Industrial Application & Impact: We provide robust, scalable aigc detection solutions for Ant Group’s diverse content platforms, including Lingguang, Jingtan, and many others.

🤝 Collaborators

We are honored to collaborate with esteemed researchers and scholars in the field of AI and Computer Vision. We deeply value these academic partnerships that drive our innovation:

  • Prof. Jun Wan (万军) | CASIA & UCAS
    • Research Interests: Biometrics, Face Anti-spoofing, Gesture Recognition, and Computer Vision.
    • [Homepage]
  • Prof. Jianfu Zhang (张健夫) | Shanghai Jiao Tong University
    • Research Interests: Computer Vision, Pattern Recognition, and Image/Video Analysis & Synthesis.
    • [Homepage]
  • Prof. Zhuosheng Zhang (张倬胜) | Shanghai Jiao Tong University
    • Research Interests: Natural Language Processing, Large Language Models, and Multi-modal Learning.
    • [Homepage]

📝 Academic Publications

  • Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection | arXiv, 2026
    • Highlights: Strengthened fine-grained visual perception for explainable AIGI detection through perception-oriented learning and value-aware on-policy distillation.
    • [Paper] [Code]
  • VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning | ICML'26, 2026
    • Highlights: Detected AI-generated videos using perception pretext reinforcement learning to capture temporal inconsistencies.
    • [Paper] [Code]
  • Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images | CVPR'26, 2026
    • Highlights: Improved detection accuracy through a two-stage approach of localizing suspicious regions followed by detailed examination.
    • [Code]
  • GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection | ICASSP'26, 2026
    • Highlights: Enhanced generalization through multi-task learning and manipulation-augmented training strategies.
    • [Paper]
  • FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning | ICLR'26, 2026
    • Highlights: Detected AI-generated images through human-aligned grounded reasoning, providing interpretable visual evidence.
    • [Paper] [Code]
  • Veritas: Generalizable deepfake detection via pattern-aware reasoning | ICLR'26 Oral, 2026
    • Highlights: Achieved generalizable deepfake detection through pattern-aware reasoning, improving robustness across diverse manipulation types.
    • [Paper] [Code]
  • Generalizable and Adaptive Continual Learning Framework for AI-generated Image Detection | IEEE TMM, 2025
    • Highlights: Proposed a continual learning framework that adapts to new generative models while mitigating catastrophic forgetting.
    • [Paper]
  • Towards explainable fake image detection with multi-modal large language models | ACM MM'25, 2025
    • Highlights: Leveraged multi-modal large language models to provide human-interpretable explanations for fake image detection.
    • [Paper]
  • WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection | AAAI'25 Oral, 2024
    • Highlights: Introduced the largest and most comprehensive AIGC image dataset at the time, providing a challenging benchmark for detection models.
    • [Paper]

🏆 Competition Achievements

  • 1st Place Winner | NTIRE 2026 Robust AI-Generated Image Detection in the Wild Challenge
    • Secured the top rank in ROC AUC for delivering superior performance in large-scale, real-world AI-generated image detection.
    • [Challenge Website]
  • 1st Place Winner | ICCV 2025 VQualA Challenge - Image Super-Resolution Generated Content Quality Assessment, 2025
    • Achieved top performance in the VQualA 2025 challenge focused on assessing the quality of super-resolution generated content.
    • [Paper 1] [Paper 2]
  • 1st Place Winner | CVPR 2024 Face Anti-Spoofing Challenge, 2024
    • Secured first place in the prestigious Face Anti-Spoofing Challenge at CVPR 2024, demonstrating state-of-the-art detection capabilities.
    • [Challenge Website]

🛠️ Open-Source Resources

  • WildFake - A large and comprehensive AIGC image detection dataset.
  • GenVideo - A large and comprehensive AIGC video detection dataset.
  • HydraFake - A large-scale challenging dataset for AI-generated image detection.
  • MintVid - A comprehensive video dataset for AIGC detection research.

✉️ Contact Us

For questions or collaborations, please contact:

⬆ Back to Top


Star History

Star History

Star History Chart
aigc-detection
ai-generated-detection
ai-generated-image-detection
awesome-list
multimodal-large-language-models

Contributors

YuZijian

26 commits

EricTan7

22 commits

yelanlan

3 commits

ant-research/Awesome-AIGC-Image-Video-Detection

A curated collection of the latest research and resources on AI-Generated Image and Video Detection.

See the code

README

Awesome AIGC Image/Video Detection Awesome

Overview

A curated collection of the latest research and resources on AI-Generated Image and Video Detection. This repository encompasses datasets, benchmarks, research papers, and practical detection tools.

🚀🚀🚀Contributions are welcome! If you find any missing papers, datasets, or tools, feel free to open an issue or submit a pull request.

Contents


🔥 Hot Events


Benchmarks & Datasets

Modality Legend: [I] Image | [V] Video | [M] Multi-modal

Annotation Type Legend: Au: Authenticity | Ex: Explainability | Lo: Localization

BenchmarkPaperVenue & YearModalityNotesReal SourceFake Source/GeneratorAnnotationScaleDownload
DF26DF26: We Cannot Tell Fake From Real AnymoreArxiv 2026[V]Modern-Generator AIGV Benchmark, Single-Person Public-Speaking Scenarios (Direct-to-Camera / Official Statement / Studio Interview), per-clip scene-prompt control (identity & scene fixed, generator varied), human + SOTA detector study near random chance, closed-source subset evaluation-onlyOpenVid-1M, TalkingCelebs, MAVOS-DD7 recent T2V/I2V models: Wan2.6, Veo 3.1, Grok Imagine 1.0, Kling 3.0 (closed); Wan2.2-A14B, HunyuanVideo 1.5, LTX 2.3 (open)Au2.7K (271 real + 2,420 fake, 1280×720)DF26 (controlled access)
DailyBenchDailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative ModelsArxiv 2026 (v3 2026-09-01)[I]Unified AIID testbed = FakeBench (full T2I synthesis) + ManipulationBench (object-level edits on real images); LAION-Aesthetics V2 pool filtered by aesthetic score ≥6 & shortest side ≥512; per-generator real/fake pairs; robustness study under recompression & pixel perturbation; ships FPD diagnostic baselineLAION-Aesthetics V2T2I: SD3.5-Large, FLUX.1, FLUX.2, Qwen-Image-2512, Z-Image, Nano Banana 2, GPT-Image 2; Edit: FLUX-Fill (random/object mask), FLUX.2-klein-9B, Qwen-Image-Edit-2511, Step1X-Edit-v1p2Au270K(v3:FakeBench ≈185K + ManipulationBench ≈75K)DailyBench
Project
AGIDefect-4KAGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and ExplanationACM MM 2026[I]Defect Detection, Localization & Explanation, 15 SOTA Generators, Quality ScoringDALL-E 3, Midjourney, FLUX, Gemini, GPT-Image, Ideogram, Kling, Grok, etc.Au, Lo, Ex4KAGIDefect-4K
RA-BenchCan We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social DisseminationArxiv 2026[V]Real-Crisis-Anchored, Source-Matched Evaluation, Human-Proof Subset, Propagation RobustnessReal crisis event footage (public media & source URLs)4 open-source + 5 closed-source generators (incl. Wan2.2)Au17.9KRA-Bench
RealHDRealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated ImagesArxiv 2026[I]Multi-category, SOTA Generators, 10K+ Prompts, Inpainting MasksT2I, Inpainting, Refinement, Face SwappingAu, Lo730K+RealHD
TreasureFleet: Few Shots Lead Effective AI-generated Image DetectionICML 2026[I]64 Models, 20 Closed-source Commercial Engines, Few-shot AdaptationDiverse architectures & 20 commercial enginesAu360KTreasure
LADBenchLADBench: A Benchmark for Logical Fault Detection in ImagesICDL 2026[I]Logical Anomaly Detection, VLM Evaluation, Common Sense ReasoningSynthetic images with logical anomaliesAu1K+LADBench
EVID-BenchWhen Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation DetectionArxiv 2026[V]Search-Grounded Verification, Evidence-Dependent ManipulationAuEVID-Bench
CoCoVideoCoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video DetectionArxiv 2026[V]Commercial AIGC Models, Contrastive BenchmarkCommercial video generation modelsAuCoCoVideo
FraudBenchFraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund EvidenceArxiv 2026[M]Fraudulent Refund Detection (ECommerce)AuFraudBench
GPT-Image-2 WildGPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of DeploymentArxiv 2026[I]GPT-Image-2Twitter (real images)GPT-Image-2Au10KGPT-Image-2 Wild
Artifact-BenchArtifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated VideosArxiv 2026[V]MLLM Evaluation, Video ArtifactsAuArtifact-Bench
CommGen15PGC: Peak-Guided Calibration for Generalizable AI-Generated Image DetectionICML 2026[I]15 Commercial Generative Models15 commercial generatorsAuCommGen15
AEGIS-AcademicAEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic ImagesArxiv 2026[I]Academic Image ForensicsAuAEGIS
SciFigDetectSciFigDetect: A Benchmark for AI-Generated Scientific Figure DetectionArxiv 2026[I]Scientific Figure DetectionNano Banana Pro, GPT-image-1.5Au150KSciFigDetect
ActivityForensicsActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosCVPR 2026[V]Action-level AIGC in videosAu6KActivityForensics
MintVidVideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningICML 2026[V]OpenVid, VFHQ, HDTF, TikTokJimeng3.0-Pro, Seedance, Kling2.5-Turbo, Sora2, TikTok, Youtube, etc.Au4KMintVid
AIGVDBenchYour One-Stop Solution for AI-Generated Video DetectionCVPR 2026[V]OpenVid-HD31 generation modelsAu440kAIGVDBench
HydraFakeVeritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningICLR 2026(Oral)[I]FFHQ, VFHQ, CelebAHQ, FF++, etc.GPT-4o, HailuoAI, ICLight, InfiniteYou, etc.Au, Ex100KHydraFake
BR-GenZooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification ApproachAAAI 2026[I]Au, Lo150KBR-Gen
RRDatasetBridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging ScenariosICCV 2025[I]Real-World Robustness, Internet Transmission, Re-digitizationAuDownLoad
HiResolutionNo Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image DetectionICLR 2026[I]Au50KHiRes-50K
AIGI-NowAlignGemini: Generalizable AI-Generated Image Detection Through Task-Model AlignmentArxiv 2026[I]COCONano Banana, GPT-4o, Jimeng, Kling, Minimax, etc.Au18KAIGI-Now
RealChainBeyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image DetectionArxiv 2026[I]Au14KRealChain
GenVidBenchGenVidBench: A 6-Million Benchmark for AI-Generated Video DetectionAAAI 2026[V]Au6MGenVidBench
SkyraSkyra: AI-Generated Video Detection via Grounded Artifact ReasoningCVPR 2026[V]Au, Ex, Lo4KViF-CoT-4K
So-Fake-SetSo-Fake: Benchmarking and Explaining Social Media Image Forgery DetectionArxiv 2025[I]F30k, WIDER, FFHQ, CelebA, OpenImages, COCO, OpenForensicsQwen-image, GPT-4o, Nano Banana, Seedream3.0, Ideogram3.0, etc.Au2M+So-Fake-Set
So-Fake-OOD
GenBuster++BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLMArxiv 2025[M]Au4KGenBuster++
GenBusterBusterX: MLLM-Powered AI-Generated Video Forgery Detection and ExplanationArxiv 2025[I]Au200KGenBuster-200K
AIGIBenchIs Artificial Intelligence Generated Image Detection a Solved Problem?NeurIPS 2025[I]FFHQ, CelebA-HQ, Open Images V7Common generators & SocialRF, CommunityAIAu200KAIGIBench
Ivy-FakeIVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC DetectionArxiv 2025[M]Au, Ex150KIvy-Fake
AEGISAEGIS: Authenticity Evaluation Benchmark for AI-Generated Video SequencesACM MM 2025[V]Vript (YouTube, TikTok), DVF, YouTube (self-collected)Stable Video Diffusion, CogVideoX-5B, I2VGen-XL, Pika, KLing, SoraAu, Ex10K+AEGIS
NeXT-IMDLNeXT-IMDL: Build Benchmark for Next-Generation Image Manipulation Detection & LocalizationArxiv 2025[I]Flickr30k, COCO, OpenImages V7SD2-Inpainting, SDXL-Inpainting, FLUX-Inpainting, etc.Au, Lo558KNeXT-IMDL
ARForensicsD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionICCV 2025[I]ImageNetInfinity, Janus_Pro, RAR, Switti, VAR, LlamaGen, Open_MAGVIT2Au300kARForensics
OpenSDIOpenSDI: Spotting Diffusion-Generated Images in the Open WorldCVPR 2025[I]Megalith-10MSD1.5, SD2.1, SDXL, SD3, Flux.1Au, Lo300KOpenSDI
Community ForensicsCommunity Forensics: Using Thousands of Generators to Train Fake Image DetectorsCVPR 2025[I]LAION, ImageNet, COCO, FFHQ, CelebA, MetFaces, AFHQ, etc.4803 generators (Latent Diffusion, GAN, Autoregressive, Pixel Diffusion, Commercial)Au2.7MCommunity Forensics
FakeClueSpot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationNeurIPS 2025[I]Au, Ex100KFakeClue
XAIGID-RewardBenchExplainable AI-Generated Image Detection RewardBenchNeurIPS 2025 Workshop[I]COCO-2017Imagen 4, Flux.1 Dev, Bagel, etc.Au, Ex3KXAIGID-RewardBench
RewardDataLearning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMsArxiv 2025[V]Au, Ex4.3KRewardData
OpenFakeOPENFAKE: An Open Dataset and Platform Toward Real-World Deepfake DetectionArxiv 2025[I]LAION-400MSD 1.5/2.1/XL/3.5, Flux 1.0-dev/1.1-Pro/Schnell, Midjourney v6/v7, DALL·E 3, Imagen 3/4, GPT Image 1, Ideogram 3.0, Grok-2, HiDream-I1, Recraft v3, Chroma, and 10 community LoRA/finetune variantsAu~4MOPENFAKE
Video Reality TestVideo Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?Arxiv 2025[V]YouTube ASMR (social media)Veo3.1-Fast, Sora2, Wan2.2-A14B, Wan2.2-5B, OpenSora-V2, HunyuanVideo, StepVideoAu149 real + dynamic fakeVideo Reality Test
DDLDDL: A Dataset for Interpretable Deepfake Detection and Localization in Real-World ScenariosArxiv 2025[M]Au367KDDL
DiffSeg30kDiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC DetectionArxiv 2025[I]COCOSD2, SD3.5, SDXL, Flux.1, Glide, Kolors, HunyuanDiT1.1, Kandinsky 2.2Au, Lo30KDiffSeg30k
FakePartsFakeParts: a New Family of AI-Generated DeepFakesArxiv 2025[V]Au, Lo81KFakeParts
ForensicHubForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and LocalizationNeurIPS 2025[I]ProGAN, StyleGAN, LDM, SDv1.4, SDv1.5, SDv2, SDXL, SD-ControlNet, MidJourney, ADM, GLIDE, VQDM, BigGANAu, Lo23 datasets
42 models
ForensicHub
LOKILOKI: A Comprehensive Synthetic Data Detection Benchmark Using Large Multimodal ModelsICLR 2025[M]SORA, Keling, Open-Sora, FLUX, Midjourney, Stable Diffusion, Nerf-based, Gaussian-based, GPT-4o, Qwen-Max, Llama 3.1-405B, MusicGen, AudioLDM2...Au, Ex18KLOKI
ChameleonA Sanity Check for AI-Generated Image DetectionICLR 2025[I]UnsplashMidjourney, DALLE-3, Stable Diffusion (various LoRA fine-tuned)Au26KChameleon
WildFakeWildFake: A Large-scale Challenging Dataset for AI-Generated Images DetectionAAAI 2025[I]Au3.7MWildFake
WildRFReal-Time Deepfake Detection in the Real-WorldArxiv 2024[I]Reddit, X (Twitter), Facebook (real images)Reddit, X (Twitter), Facebook (social media deepfakes)AuWidlRF
AIGCDetectBenchmarkPatchCraft: Exploring Texture Patch for Efficient AI-generated Image DetectionArxiv 2024[I]Au100KAIGCDetectionBenchMark
GenVideoDeMamba: AI-Generated Video Detection on Million-Scale GenVideo BenchmarkArxiv 2024[V]Au2.3MGenVideo
DRCTDrct: Diffusion reconstruction contrastive training towards universal detection of diffusion generated imagesICML 2024[I]MSCOCOLDM, SDv1.4, SDv1.5, SDv2, SDXL, SD-ControlNetAu2MDRCT-2M
GenImageGenImage: A Million-Scale Benchmark for Detecting AI-Generated ImageNeurIPS 2023[I]ImageNet, WukongMidJourney, SDv1.4, SDv1.5, ADM, GLIDE, VQDM, BigGANAu2.7MGenImage
DF40DF40: Toward Next-Generation Deepfake DetectionNeurIPS 2024[I] [V]Au0.1M+ videos, 1M+ imagesDF40
Forensics-BenchForensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language ModelsCVPR 2025[I], [V], [M]Various public datasetsGAN, Diffusion, VAE, RNN, Encoder-Decoder, Graphics-basedAu, Lo63KForensics-Bench

⬆ Back to Top


Research Papers

💡 Note: Papers are sorted by year (descending) within each category.
Modality Legend: [I] Image | [V] Video | [M] Multi-modal
Category Layout: Top level is split into MLLM-Based (MLLM-powered detection) and Classification-Based (compact/traditional classifiers). Classification-Based is further organized into six subcategories. When a paper fits multiple subcategories, the priority is: Training-Free/Zero-Shot > Continual/Incremental > Video Spatiotemporal > Frequency/Low-Level Artifacts > Supervised General.

MLLM-Based

This category focuses on utilizing Multimodal Large Language Models (MLLMs) like GPT-4V, LLaVA, or Qwen-VL to detect AI-generated content. These methods often provide natural language explanations (explainability) alongside binary detection.

TitleVenue & YearModalityHighlights/KeywordsCode
Evidence-Guided Detection, Localization and Explanation for Text-Centric Image ForensicsACM MM 2026 Challenge[I]Detector-Localizer-Reasoner Cascade, Iterative Difficulty-Aware Mining, Report-Mask ConsistencyGitHub
AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and ExplanationACM MM 2026[I]AGIDefect-4K Dataset, Hierarchical Defect Annotation, MLLM Baseline (AGIDA)GitHub
Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation OptimizationACM MM 2026[I]Feature-robust Augmentation, Mean-Teacher Consistency, Evidence-grounded Preference Optimization, Challenge WinnerGitHub
PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMsIJCAI 2026 Workshop[I]Perception-as-Tool, DINOv3 Forensic Perception Tool, General-Purpose MLLM ExplanationGitHub
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI DetectionACM MM 2026[I]Interactive Visual Search, Verifier-guided Evidence Alignment, GroundFake Dataset, FakeFrontier BenchmarkN/A
SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited DataArxiv 2026[I]Adversarial RL Loop, Diffusion Editor, Free-form ExplanationN/A
VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsArxiv 2026[V]Meta-Detection RL, Verifiable Temporal Grounding, Evidence-Guided Reward RedistributionN/A
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI DetectionArxiv 2026[I]Value-aware On-Policy Distillation, Perception-Enhanced ReasoningGitHub
Detecting AI-Generated Video: A Vision-Language Dual-View SurveyACL 2026 Findings[V][Survey] Vision-Language Dual-View Taxonomy, Factual Fidelity Verification, Cross-modal ConsistencyN/A
TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image DetectionICML 2026[I]Artifact feature and Semantic feature fusionGitHub
Venus-DeFakerOne: Unified Fake Image Detection & LocalizationArxiv 2026[I]Unified Detection & Localization, Large-Scale TrainingGitHub
GenShield: Unified Detection and Artifact Correction for AI-Generated ImagesICML 2026[I]Unified MLLM, Detect & Correct ArtifactsGitHub
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationArxiv 2026[I]Reasoning-Aligned Representation, InterpretableN/A
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image DetectionCVPR 2026[I]Unified Model, Co-Evolution (Generation & Detection)GitHub
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement LearningICML 2026[V]Perception Pretext RL, Fact-based Reasoning, MintVid DatasetGitHub
Veritas: Generalizable deepfake detection via pattern-aware reasoningICLR 2026(Oral)[I]Pattern-aware Reasoning, HydraFake DatasetGithub
DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-ReflectionArxiv 2026[I]Knowledge Injection, Self-ReflectionN/A
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic ReasoningArxiv 2026[M]Agentic Framework, Document SafetyN/A
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RLICLR 2026[V]Multi-stage RL (GRPO), Time Artifacts, Video Detection DatasetGitHub
FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningICLR 2026[I]Grounded Reasoning, Human-annotated DatasetN/A
AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model AlignmentArxiv 2026[I]Decoupling (Semantic & Pixel), AIGI-Now DatasetN/A
Zoom-In to Sort AI-Generated Images OutArxiv 2026[I]Thinking with Images, MagniFake DatasetN/A
AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image DetectionArxiv 2026[I]Agentic frameworkGithub
EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image DetectionArxiv 2026[I]Agentic Framework, Method EnsemblingN/A
VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake DetectionArxiv 2026[I]Part-centric Forensic, OmniFake DatasetProject
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?Arxiv 2026[V]GenVideoLens benchmarkN/A
Semantic Visual Anomaly Detection and Reasoning in AI-Generated ImagesICLR 2026[I]Semantic Anomaly Reasoning, AnomReason DatasetN/A
FAKE-HR1: RETHINKING REASONING OF VISION LANGUAGE MODEL FOR SYNTHETIC IMAGE DETECTIONArxiv 2026[I]Hybrid-Reasoning, Dual-mode DatasetN/A
MIRAGE: Towards AI-Generated Image Detection in the WildArxiv 2025[I]Human Curation Dataset, Heuristic-to-Analytic ReasoningN/A
BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLMArxiv 2025[M]RL Post-training, Cross-Modal, Thinking Reward MechanismGithub
BusterX: MLLM-Powered AI-Generated Video Forgery Detection and ExplanationArxiv 2025[V]GenBuster-200K Dataset, Cold Start + RL TrainingGithub
REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image DetectionArxiv 2025[I]Chain-of-Evidence, Expert-grounded RLN/A
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationNeurIPS 2025[I]FakeClue Dataset, Fine-grained Artifact Clues, Artifact ExplanationGitHub
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsICCV 2025[I]Holmes-Set, Multi-Expert Jury, 3-Stage Training PipelineGithub
LEGION: Learning to Ground and Explain for Synthetic Image DetectionICCV 2025[I]SynthScars Dataset, Defender & Controller, Image RefinementGitHub
Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image DetectionArxiv 2025[I]Perception & Reasoning, ExplainFake-BenchN/A
SIDA: Social Media Image Deepfake Detection, Localization, and ExplanationCVPR 2025[I]SID-Set, Mask Prediction, Social Media ContextGithub
FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language ModelsICLR 2025[I]Explainable IFDL, Domain Tag-guided, Multi-modal LocalizationGitHub
FakeScope: Large Multimodal Expert Model for Transparent AI-Generated Image ForensicsArxiv 2025[I]FakeChain Dataset, FakeInstruct, Trace EvidenceN/A
AntifakePrompt: Prompt-Tuned Vision-Language Models are Fake Image DetectorsArxiv 2024[I]VQA, InstructBLIP, Soft Prompt-tuning, Zero-shotGitHub

Classification-Based

This category includes supervised learning approaches that train neural networks (CNNs, ViTs, VFMs, etc.) specifically to classify authentic vs. AI-generated content. They usually focus on robustness, generalization, and feature extraction. It is organized into six subcategories: Supervised General Detectors, Video Spatiotemporal Modeling, Training-Free / Zero-Shot, Frequency-Domain & Low-Level Artifacts, Continual & Incremental Learning, and Related & Other.

Supervised General Detectors

Trainable classifiers and backbones (CNNs, ViTs, vision foundation models, CLIP-based adapters, few-shot and prompt-based methods) for general-purpose detection.

TitleVenue & YearModalityHighlights/KeywordsCode
Learning Continuous Source Responses For Generalizable AI-Generated Image DetectionArxiv 2026[I]CuRe, Continuous Source-Response Regression (Real-Generated Mixing Ratio), Shortcut-Cue Suppression, Source-Response Subspace, 10-Benchmark EvaluationGitHub
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image DetectionArxiv 2026[I]UCF-Net, CLIP Semantic + DINO Structural Priors, Layer-wise Expert Aggregation, Entropy-based Uncertainty Fusion, 4M-image Unified Benchmark, Cross-generator EvaluationGitHub
FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and LocalizationArxiv 2026[I]Forensic-Semantic MoE, Joint Detection & Localization, Cross-generator Generalization (OpenSDID)GitHub
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation LocalizationArxiv 2026[I]Global Artifact Token, SAM3 FiLM Injection, Boundary Adhesion AnalysisN/A
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in Semantic ResidualsECCV 2026 (Spotlight)[I]Low-Rank Collapse Signature, Semantic-Residual Decoupling, Cross-model GeneralizationN/A
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image DetectionECCV 2026[I]Closed-form Gaussian Heads, Percept-Lens Protocol (39 Datasets), Transfer DiagnosticN/A
Environment-Invariant Subspace Learning for Generalizable Deepfake DetectionArxiv 2026[I]Environment-Invariant Subspace, VFM Semantic Priors, Environmental InterventionN/A
Understanding Why Foundation Models Work for Diffusion-Generated Image DetectionArxiv 2026[I]Interpretability Analysis, DDIM Inversion, Low-to-Mid Frequency Distributional DiscrepancyN/A
PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image DetectionArxiv 2026[I]DINO Patch Tokens, 2D Spatial Aggregation, LoRA AdaptersN/A
GlobalForge: Towards Robust AI-Generated Image DetectionArxiv 2026[I]Global Structural Reasoning, Local Information Bottleneck, RealDeg-BenchCode
Fleet: Few Shots Lead Effective AI-generated Image DetectionICML 2026[I]Few-shot Adaptation, Routing Correction, Treasure BenchmarkGitHub
SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision EncodersArxiv 2026[I]Frozen Vision Encoders, Linear Classifier, RealWorldBenchN/A
HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image DetectionACM MM 2026[I]Asymmetric Prompting, Dynamic Decision BoundaryN/A
VINA: Video as Natural Augmentation: Towards Unified AI-Generated Image and Video DetectionArxiv 2026[M]Unified Image/Video Detection, Cross-Modal Contrastive LearningN/A
PGC: Peak-Guided Calibration for Generalizable AI-Generated Image DetectionICML 2026[I]Peak-Guided Calibration, CommGen15 DatasetGitHub
Reduce the Artifacts Bias for More Generalizable AI-Generated Image DetectionArxiv 2026[I]Bias-free TrainingGitHub
Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification ApproachAAAI 2026[I]Localized AIGC Detection, Forgery Amplification, Scene-aware Local ForgeryGitHub
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation ModelsArxiv 2026[I]Linear Probe, Vision Foundation Models, Emergent Forensic CapabilityN/A
MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image DetectionArxiv 2026[I]Manifold Reconstruction, Memory Bank, Human-AIGI BenchmarkGitHub
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image DetectionICLR 2026[I]Detail-preserving dual-path architecture, Multi-task learning, HiRes-50K benchmarkN/A
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch LearningICLR 2026[I]Random Patch Replacement, Patch-wise Contrastive LearningN/A
OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the WildArxiv 2025[I]Mixture-of-Experts, Semantic-Artifact Decoupling, Mirage DatasetGitHub
DINO-Detect: A Simple yet Effective Framework for Blur-Robust AI-Generated Image DetectionArxiv 2025[I]Blur Robustness, Knowledge Distillation, DINOv3Github
Orthogonal Subspace Decomposition for Generalizable AI-Generated Image DetectionICML 2025 (Oral)[I]SVD Orthogonal Subspace, Asymmetry Phenomenon, Parameter-efficient Fine-tuningGitHub
A Bias-Free Training Paradigm for More General AI-generated Image DetectionCVPR 2025[I]Bias-Free, Semantic Alignment, Stable Diffusion Self-conditioningGithub
Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionCVPR 2025[I]CLIP, Blending Boundaries, Forgery-aware Prompt LearningGithub
Exploring Unbiased Deepfake Detection via Token-Level Shuffling and MixingAAAI 2025[I]Token-Level Shuffling, Contrastive Loss, Bias MitigationN/A
FakeFormer: Efficient Vulnerability-Driven Transformers for Generalisable Deepfake DetectionArxiv 2024[I]Vulnerability-driven, Local Attention (L2-Att), Vision TransformerGitHub

Video Spatiotemporal Modeling

Methods that exploit temporal inconsistencies, motion patterns, and spatiotemporal artifacts in AI-generated videos.

TitleVenue & YearModalityHighlights/KeywordsCode
Mind the Rift: Cross-Scale Coupling Mismatch for AI-Generated Video DetectionACM MM 2026[V]Cross-Scale Coupling Mismatch, Macro Temporal Dynamics vs. Pixel-Level Residuals, Persistent Homology, Encoder-AgnosticGitHub
MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow TrajectoriesArxiv 2026[V]Physical Motion Consistency, Sparse Optical-Flow Trajectories, Multi-scale Geometric EvolutionN/A
Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video DetectionArxiv 2026[V]V-PVP Readout, Patch Velocity Profiling, Frozen Video BackbonesCode
Dataset Biases and Shortcut Learning in Motion-Based AI-Generated Video DetectionArxiv 2026[V]Motion Bias Analysis, Preprocessing/Sampling Bias, Frequency-based ComparisonN/A
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal DisentanglementArxiv 2026[V]Counterfactual Intervention, Causal Disentanglement, Cross-domain GeneralizationGitHub
ReConFuse: Reconstruction-Error Guided Semantic Fusion for AI-Generated Video DetectionArxiv 2026[V]Reconstruction Error, Semantic Fusion, Spatial-Temporal ArtifactsN/A
Detecting AI-Generated Videos with Spiking Neural NetworksArxiv 2026[V]Spiking Neural Networks, Temporal ArtifactN/A
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video DetectionArxiv 2026[V]Cross-Modal Temporal Artifacts, Video DetectionN/A
Preserving Forgery Artifacts: AI-Generated Video Detection at Native ScaleICLR 2026[V]Native scale video processing, Massive realistic video dataset, Preserves subtle generation artifactsN/A
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationNeurIPS 2025[V]Wavelet-band Augmentation, Forensic Frequency Artifacts, Single-generator GeneralizationGitHub
AI-Generated Video Detection via Perceptual StraighteningNeurIPS 2025[V]Perceptual Straightening, DINOv2, Temporal CurvatureGitHub
Physics-Driven Spatiotemporal Modeling for AI-Generated Video DetectionNeurIPS 2025[V]Normalized Spatiotemporal Gradient (NSG), Maximum Mean Discrepancy (MMD)Github
Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated ContentCVPR 2025[V]SigLIP-So400M, Attention-Diversity Loss, Full-frame ManipulationsN/A
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake DetectionTMM 2025[V]Direction-aware Attention, SpatioTemporal Invariant LossN/A
DeMamba: AI-Generated Video Detection on Million-Scale GenVideo BenchmarkArxiv 2024[V]Mamba, State Space Model, Long-range Spatiotemporal InconsistencyGitHub
Distinguish Any Fake Videos: Unleashing the Power of Large-scale Data and Motion FeaturesArxiv 2024[V]GenVidDet, Optical Flow, Dual-Branch 3D TransformerN/A

Training-Free / Zero-Shot

Methods that detect AI-generated content without additional training on detection data.

TitleVenue & YearModalityHighlights/KeywordsCode
Frozen DINO Localizes Image Edits Without a LocalizerArxiv 2026[I]Training-free, Frozen DINO Patch-token Drift, Haar Perturbation, Edit LocalizationGitHub
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal RoughnessECCV 2026[V]Training-free, Patch-level Incoherence, Temporal Roughness, Ultra-low FPRGitHub
Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsCVPR 2026[V]Training-free, Zero-shot, Spatial-Temporal Likelihoods, ComGenVid DatasetGitHub

Frequency-Domain & Low-Level Artifacts

Methods based on spectral analysis, quantization/upsampling traces, and other low-level generative artifacts.

TitleVenue & YearModalityHighlights/KeywordsCode
Structured Local Differential Modeling for AI-Generated Image DetectionArxiv 2026[I]RippleNet, Local Differential Signals, Low-SNR Forgery TracesN/A
Dual Data Alignment Makes AI-Generated Image Detector Easier GeneralizableNeurIPS 2025 (Spotlight)[I]Dual-domain Alignment, Frequency-level Bias, VAE ReconstructionGitHub
D3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image DetectionICCV 2025[I]Discrete Distribution Discrepancy-aware Transformer, Vector Quantized Variational AutoEncoderGithub
Any-Resolution AI-Generated Image Detection by Spectral LearningCVPR 2025[I]Spectral Context Attention, Frequency Reconstruction, OOD DetectionGithub
Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningAAAI 2024[I]Frequency Domain, FFT, Frequency Conv Layer (FCL), LightweightGitHub
Rethinking the Up-Sampling Operations in CNN-based Generative NetworkCVPR 2024[I]Neighboring Pixel Relationships, Generalized Structural ArtifactsGithub

Continual & Incremental Learning

Methods that keep adapting detectors to evolving generators without catastrophic forgetting.

TitleVenue & YearModalityHighlights/KeywordsCode
Preserving Knowledge across Space and Time for Continual Video Deepfake DetectionECCV 2026[V]Modality-Specific Frequency Distillation, Spatial/Temporal/Spatiotemporal Decomposition, Cross-Modality DecorrelationGitHub
Automated In-the-Wild Data Collection for Continual AI Generated Image DetectionArxiv 2026[I]Continual Learning, Continual Data CollectionGitHub
IncreFA: Breaking the Static Wall of Generative Model AttributionArxiv 2026[I]Incremental Learning, Generative Model AttributionGitHub
SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual LearningCVPR 2026[I]Scene-aware optimization, Continual learningGitHub
Generalizable and Adaptive Continual Learning Framework for AI-generated Image DetectionTMM 2026[I]Continual Learning, Kronecker-Factored Approximate CurvatureN/A

Papers related to AI-generated content safety (e.g., provenance/watermarking, misinformation verification) that do not fit the subcategories above.

TitleVenue & YearModalityHighlights/KeywordsCode
DF26: We Cannot Tell Fake From Real AnymoreArxiv 2026[V][Benchmark] 2,691 Public-Speaking Videos, 7 Modern T2V/I2V Models, Human & SOTA Detectors Near Chance, Distribution-Shift RobustnessN/A
APT: Anchor-aligned Perturbations for Tamper Localization in Fully Regenerated ImagesECCV 2026[I][Proactive Forensics] Semi-Fragile Latent Perturbation, Fully Regenerated (Inpainting) Setting, Anchor-Direction AlignmentN/A
Training-Free Reconstruction-Based AI-Generated Image Detectors Are Inherently Vulnerable to Adversarial ExamplesECCV 2026 Workshop[I][Robustness Analysis] Reconstruction-based Detector Attacks, Transferable Adversarial Examples, Real-world DegradationsN/A
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social DisseminationArxiv 2026[V][Evaluation] RA-Bench, Crisis Event Videos, Detector Generalization, Social DisseminationN/A
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation DetectionArxiv 2026[V]Search-Grounded Verification, EVID-Bench, Evidence-Dependent ManipulationN/A
Robust ASIC-Based Image Authentication Using Reed-Solomon LSB WatermarkingPreprint 2026[I]ASIC PoW, Hardware-bound Provenance, Reed-Solomon WatermarkingGitHub

⬆ Back to Top


Competitions

CompetitionLinkYearInfo
Robust AIGC DetectionNTIRE 2026 Robust AI-Generated Image Detection in the Wild2026No restrictions on training data.
Evaluate ROC AUC metrics on robust samples.
Robust Deepfake DetectionNTIRE 2026 Robust Deepfake Detection Challenge2026No restrictions on training data.
The 6th Face Anti-spoofing ChallengeThe 6th Face Anti-Spoofing: Unified Physical-Digital Attacks Detection@ICCV20252025No external data or pre-trained models allowed.
Limited to a single DL model with under 100G FLOPs.
Detect AI vs. Human-Generated Images2025 Women in AI (WAI) Kaggle Challenge2025Paired dataset of authentic and AI-generated images
The 5th Face Anti-spoofing Challenge5th Chalearn Face Anti-spoofing Workshop and Challenge@CVPR20242024UniAttackData+ for unified physical and digital attack detection.

⬆ Back to Top


Practical Detection Tools

  • 美亚鉴真 - 微信小程序搜索 美亚鉴真
  • SiliconSignature - GitHub - Hardware-bound image authentication using ASIC PoW nonces for unforgeable provenance certification
  • EyeSift - Website - Free online AI text/image/video/audio detector with detailed per-model benchmarks
  • Hive Moderation - Website
  • Tencent Zhuque AI Detection Assistant - Website
  • AI or Not - Website
  • Illuminarty - Website
  • Winston AI - Website
  • Is it AI? - Website
  • TruthScan - Website
  • 中科睿鉴 (Zhongke Ruijian) - 微信小程序搜索 睿鉴AI

⬆ Back to Top


🏢 About Our Team

We are the Content Security Intelligence Team under Ant Group - Machine Intelligence. We are responsible for developing comprehensive content security and risk-mitigation capabilities for the Ant Group ecosystem, bridging the gap between rapidly evolving technologies and the urgent need for digital trust.

Why We Do It

In an era where synthetic media is increasingly sophisticated and pervasive, our research serves as a critical line of defense. By advancing AIGC detection technologies, we aim to:

  • Safeguard Digital Integrity: We provide essential defense mechanisms to protect the authenticity of visual content and combat the spread of misinformation in the digital space.
  • Empower Trust: Our solutions ensure the public can distinguish between genuine and synthetic media, fostering a more transparent and trustworthy digital ecosystem.
  • Industrial Application & Impact: We provide robust, scalable aigc detection solutions for Ant Group’s diverse content platforms, including Lingguang, Jingtan, and many others.

🤝 Collaborators

We are honored to collaborate with esteemed researchers and scholars in the field of AI and Computer Vision. We deeply value these academic partnerships that drive our innovation:

  • Prof. Jun Wan (万军) | CASIA & UCAS
    • Research Interests: Biometrics, Face Anti-spoofing, Gesture Recognition, and Computer Vision.
    • [Homepage]
  • Prof. Jianfu Zhang (张健夫) | Shanghai Jiao Tong University
    • Research Interests: Computer Vision, Pattern Recognition, and Image/Video Analysis & Synthesis.
    • [Homepage]
  • Prof. Zhuosheng Zhang (张倬胜) | Shanghai Jiao Tong University
    • Research Interests: Natural Language Processing, Large Language Models, and Multi-modal Learning.
    • [Homepage]

📝 Academic Publications

  • Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection | arXiv, 2026
    • Highlights: Strengthened fine-grained visual perception for explainable AIGI detection through perception-oriented learning and value-aware on-policy distillation.
    • [Paper] [Code]
  • VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning | ICML'26, 2026
    • Highlights: Detected AI-generated videos using perception pretext reinforcement learning to capture temporal inconsistencies.
    • [Paper] [Code]
  • Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images | CVPR'26, 2026
    • Highlights: Improved detection accuracy through a two-stage approach of localizing suspicious regions followed by detailed examination.
    • [Code]
  • GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection | ICASSP'26, 2026
    • Highlights: Enhanced generalization through multi-task learning and manipulation-augmented training strategies.
    • [Paper]
  • FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning | ICLR'26, 2026
    • Highlights: Detected AI-generated images through human-aligned grounded reasoning, providing interpretable visual evidence.
    • [Paper] [Code]
  • Veritas: Generalizable deepfake detection via pattern-aware reasoning | ICLR'26 Oral, 2026
    • Highlights: Achieved generalizable deepfake detection through pattern-aware reasoning, improving robustness across diverse manipulation types.
    • [Paper] [Code]
  • Generalizable and Adaptive Continual Learning Framework for AI-generated Image Detection | IEEE TMM, 2025
    • Highlights: Proposed a continual learning framework that adapts to new generative models while mitigating catastrophic forgetting.
    • [Paper]
  • Towards explainable fake image detection with multi-modal large language models | ACM MM'25, 2025
    • Highlights: Leveraged multi-modal large language models to provide human-interpretable explanations for fake image detection.
    • [Paper]
  • WildFake: A Large-scale Challenging Dataset for AI-Generated Images Detection | AAAI'25 Oral, 2024
    • Highlights: Introduced the largest and most comprehensive AIGC image dataset at the time, providing a challenging benchmark for detection models.
    • [Paper]

🏆 Competition Achievements

  • 1st Place Winner | NTIRE 2026 Robust AI-Generated Image Detection in the Wild Challenge
    • Secured the top rank in ROC AUC for delivering superior performance in large-scale, real-world AI-generated image detection.
    • [Challenge Website]
  • 1st Place Winner | ICCV 2025 VQualA Challenge - Image Super-Resolution Generated Content Quality Assessment, 2025
    • Achieved top performance in the VQualA 2025 challenge focused on assessing the quality of super-resolution generated content.
    • [Paper 1] [Paper 2]
  • 1st Place Winner | CVPR 2024 Face Anti-Spoofing Challenge, 2024
    • Secured first place in the prestigious Face Anti-Spoofing Challenge at CVPR 2024, demonstrating state-of-the-art detection capabilities.
    • [Challenge Website]

🛠️ Open-Source Resources

  • WildFake - A large and comprehensive AIGC image detection dataset.
  • GenVideo - A large and comprehensive AIGC video detection dataset.
  • HydraFake - A large-scale challenging dataset for AI-generated image detection.
  • MintVid - A comprehensive video dataset for AIGC detection research.

✉️ Contact Us

For questions or collaborations, please contact:

⬆ Back to Top


Star History

Star History

Star History Chart
aigc-detection
ai-generated-detection
ai-generated-image-detection
awesome-list
multimodal-large-language-models

Contributors

YuZijian

26 commits

EricTan7

22 commits

yelanlan

3 commits