media-sec-lab/Audio-Deepfake-Detection

Research progress on speech deepfake detection: Relevant datasets aggregated from the review literature and publicly available codes

324

60 commits

updated Jun 4, 2025

See the code

README

[!IMPORTANT]

A curated collection of papers and resources on Audio Deepfake Detection (ADD).

Please refer to our survey "Research progress on speech deepfake and its detection techniques" for the detailed contents. Paper page

Please let us know if you discover any mistakes or have suggestions by emailing us: xuyuxiong2022@email.szu.edu.cn

Table of contents

What's New

Survey

⬆ Back to top

  • 【2022-04】-【ADD-Survey】-【Ben Gurion University】
    • A study on data augmentation in voice anti-spoofing
    • Author(s): Ariel Cohen, Inbal Rimon, Eran Aflalo, Haim H. Permuter
    • Paper Code
  • 【2022-05】-【ADD-Survey】-【King Saud University】
    • A review of modern audio deepfake detection methods: challenges and future directions
    • Author(s): Zaynab Almutairi, and Hebah Elgibreen
    • Paper
  • 【2023-01】-【ADD-Survey】-【University of Maryland Baltimore County】
    • Audio deepfakes: A survey
    • Author(s): Zahra Khanjani, Gabrielle Watson and Vandana P. Janeja
    • Paper
  • 【2023-02】-【ADD-Survey】-【Oakland University】
    • Battling voice spoofing: a review, comparative analysis, and generalizability evaluation of state-of-the-art voice spoofing counter measures
    • Author(s): Awais Khan, Khalid Mahmood Malik, James Ryan, and Mikul Saravanan
    • Paper Code
  • 【2023-04】-【ADD-Survey】-【Panjab University】
    • Review of audio deepfake detection techniques: Issues and prospects
    • Author(s): Abhishek Dixit, Nirmal Kaur, Staffy Kingra
    • Paper
  • 【2023-07】-【ADD-Survey】-【Indian Institute of Technology】
    • Uncovering the deceptions: an analysis on audio spoofing detection and future prospects
    • Author(s): Rishabh Ranjan, Mayank Vatsa, Richa Singh
    • Paper
  • 【2023-08】-【ADD-Survey】-【CASIA】
    • Audio deepfake detection: a survey
    • Author(s): Jiangyan Yi, Chenglong Wang, Jianhua Tao, Xiaohui Zhang, Chu Yuan Zhang, and Yan Zhao
    • Paper
  • 【2024-02】-【Multimodal-Survey】-【Purdue University, Nanchang University】
    • Detecting Multimedia Generated by Large AI Models: A Survey
    • Author(s): Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu
    • Paper Code
  • 【2024-04】-【ADD-Survey】-【Toronto Metropolitan University】
    • A Survey on Speech Deepfake Detection
    • Author(s): Menglu Li, Yasaman Ahmadiadli, Xiao-Ping Zhang
    • Paper
  • 【2024-05】-【Multimodal-Survey】-【Cool Large Language Models Research Group, Renmin University of China】
    • Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities
    • Author(s): Xiaomin Yua, Yezhaohui Wanga, Yanfang Chenb, Zhen Tao, Dinghao Xi, Shichao Song, and Simin Niu
    • Paper
  • 【2024-09】-【ADD-Survey】-【Austrian Institute of Technology】
    • A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
    • Author(s): Lam Pham, Phat Lam, Tin Nguyen, Hieu Tang, Huyen Nguyen, Alexander Schindler, Hai Canh Vu
    • Paper Code
  • 【2024-11】-【Multimodal-Survey】-【University of Bucharest】
    • Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
    • Author(s): Forinel-Alin Croitoru, Andrei-Iulian Hıˆji, Vlad Hondru, Nicolae Cat ̆alin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah
    • Paper Code
  • 【2024-11】-【Multimodal-Survey】-【University College Dublin】
    • Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey
    • Author(s): Hong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac
    • Paper
  • 【2024-12】-【ADD-Survey】-【Imperial College London, Technical University of Munich】
    • From Audio Deepfake Detection to AI-Generated Music Detection——A Pathway and Overview
    • Author(s): Yupei Li, Manuel Milling, Lucia Specia, Björn W. Schuller
    • Paper
  • 【2025-02】-【Multimodal-Survey】-【BUPT, University of California, CASIA】
    • Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
    • Author(s): Yueying Zou, Peipei Li, Zekun Li, Huaibo Huang, Xing Cui, Xuannan Liu, Chenghanyu Zhang, Ran He
    • Paper

Top Repositories

⬆ Back to top

  • Audio Large Language Models
    • Resources on Audio Large Language Models, including datasets, methods, benchmarks, and studies.
    • Repo
  • awesome-fake-audio-detection
    • A list of tools, papers and code related to Fake Audio Detection.
    • Repo
  • ASVspoof Challenge Official Repository

Audio Large Model

⬆ Back to top

ModelPublisherYearsAchievable Tasks
AudioLM
Paper Website Code
Google2022.091. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation space.
2. Speech continuation, Acoustic generation, Unconditional generation, Generation without semantic tokens, and Piano continuation.
VALL-E
Paper Website
Microsoft2023.011. Simply record a 3-second registration of an unseen speaker to create a high-quality personalised speech.
2. VALL-E X: Cross-lingual speech synthesis.
USM
Website
Google2023.031. ASR beyond 100 languages.
2. Downstream ASR tasks.
3. Automated Speech Translation (AST).
SpeechGPT
Website
Fudan University2023.051. Perceive and generate multi-modal contents.
2. Spoken dialogue LLM with strong human instruction.
Pengi
Paper Website
Microsoft2023.051. an Audio Language Model that leverages Transfer Learning by framing all audio tasks as text-generation tasks.
2. The unified architecture of Pengi enables open-ended tasks and close-ended tasks without any additional fine-tuning or task-specific extensions.
VoiceBox
Website
Meta2023.061. Synthesize speech across six languages.
2. Remove transient noise.
3. Edit content.
4. Transfer audio style within and across languages.
5. Generate diverse speech samples.
AudioPaLM
Paper Website
Google2023.061. Speech-to-speech translation.
2. Automatic Speech Recognition (ASR).

Datasets

⬆ Back to top

Attack TypesYearsDatasetNumber of Audio
(Subdataset:Real/Fake)
Language
TTS2019FoR
Paper Dataset
111,000/87,000English
TTS2021WaveFake
Paper Dataset
16,283/117,985English, Japanese
TTS2021Half-truth
Paper
53,612/107,224Chinese
TTS2021HAD
Paper
53,612/107,224Chinese
TTS2021FakeAVCeleb
Paper Dataset
10,209/11,335English
TTS2022ADD 2022
Paper
LF: 5,619/46,067
PF: 5,319/46,419
FG-D: 5,319/46,206
Chinese
TTS2022CMFD
Paper Dataset
Chinese: 1,800/1,000
English: 1,800/1,000
English, Chinese
TTS2022In-the-Wild
Paper Dataset
19,963/11,816English
TTS2022CFAD
Paper Dataset
38,600/77,200Chinese
TTS2022Psynd
Paper Dataset
2,294English
TTS2022TIMIT-TTS
Paper Dataset
0/79,120English
TTS2023ODSS
Paper Dataset
11,032/18,993English, German,
and Spanish
TTS2024MLAAD
Paper Dataset
-/76,000Multi-lingual
TTS2024CD-ADD
Paper
300 hoursEnglish
TTS2024DiffSSD
Paper Dataset
24,226/70,000English
TTS2024ACCENT
Paper
53,651/192,461Multi-lingual
TTS2024SpoofCeleb
Paper Dataset
250k+English
TTS2024LlamaPartialSpoof
Paper Dataset
10,573/33,479English
TTS2024FakeSound
Paper Dataset
-/3,798English
TTS2024DFADD
Paper Dataset
44,455/163,500English
TTS2024SONAR
Paper Dataset
-/2,274Multi-lingual
TTS20256KSFx
Paper Dataset
-/6,000-
TTS2025BangalFake
Paper Dataset
12,260/13,260Bengali
Replay2017ASVspoof 2017
Paper Dataset
3,565/14,465English
Replay2019ReMASC
Paper Dataset
9,240/45,472English, Chinese,
Hindi
Replay2019VSDC
Paper Dataset
1,687/11,772English
Replay2024POLIPHONE
Paper Dataset
41,941English
TTS
and VC
2015AVspoof
Paper Dataset
LA: 15,504/120,480
PA: 15,504/14,465
English
TTS
and VC
2015ASVspoof 2015
Paper Dataset
16,651/246,500English
TTS
and VC
2021FMFCC-A
Paper Dataset
10,000/40,000Chinese
TTS
and VC
2022SceneFake
Paper Dataset
19,838/64,642English
TTS
and VC
2022EmoFake
Paper
35,000/53,200English, Chinese
TTS
and VC
2023PartialSpoof
Paper Dataset
12,483/108,978English
TTS
and VC
2023ADD 2023
Paper
FG-D: 172,819/113,042
RL: 55,468/65,449
AR: 14,907/95,383
Chinese
TTS
and VC
2023DECRO
Paper Dataset
Chinese: 21,218/41,880
English: 12,484/42,799
English, Chinese
TTS
and VC
2023HABLA
Paper Dataset
22,000/58,000Spanish
TTS
and VC
2024DeepFakeVox-HQ
Paper
693k/643kEnglish
TTS
and VC
2024VoiceWukong
Paper Dataset
5,300/413,400English, Chinese
TTS
and VC
2024Speech-Forensics
Paper Dataset
13,100/7,452English
TTS
and VC
2024VoiceEdit
Paper
-Multi-lingual
TTS
and VC
2024RFP
Paper Dataset
28,115/74,199English
TTS
and VC
2025MADD
Paper
60,000/129,990Multi-lingual
TTS
and VC
2025XMAD-Bench
Paper Dataset
414,858Multi-lingual
TTS
and Vocoder
2024Diffuse or Confuse
Paper Dataset
131,000/183,400English
TTS
and Vocoder
2025ShiftySpeech
Paper Dataset
3,000+ hoursEnglish, Chinese,
Japanese
TTS, VC
and Replay
2019ASVspoof 2019
Paper Dataset
LA: 12,483/108,978
PA: 28,890/189,540
English
TTS, VC
and Replay
2021ASVspoof 2021
Paper Dataset
LA: 18,452/163,114
PA: 126,630/816,480
PF: 14,869/519,059
English
TTS, VC
and Vocoder
2024SpeechFake
Paper
3,000+ hoursEnglish
VC, Replay
and Adversarial
2024VSASV
Paper Dataset
164,000/174,000Multi-lingual
TTS, VC
and Adversarial
2024ASVspoof 5
Paper Dataset
188,819/ 815,262English
Voice Cloning2021RTVCSpoof
Paper
3,284/4,843English
Voice Cloning2024Kratika Dataset
Paper
24,226/25,000English
Vocoder2022Yan Dataset
Paper
8,200/63,200Chinese
Vocoder2023LibriSeVoc
Paper Dataset
13,201/79,206English
Vocoder2023Voc.v1-v4
Paper Dataset
2,580/10,320English
Vocoder2024MLADDC
Paper Dataset
80k/160kMulti-lingual
Vocoder2024CVoiceFake
Paper Dataset
23,544/91,700Multi-lingual
Impersonation2024IPAD
Paper
5,170/18,874Chinese
Text-To-Music (TTM)2024FSD
Paper Dataset
200/500Chinese
Text-To-Music (TTM)2024SingFake
Paper Dataset
634/671Multi-lingual
Text-To-Music (TTM)2024CtrSVDD
Paper Dataset
32,312/188,486Multi-lingual
Text-To-Music (TTM)2024FakeMusicCaps
Paper Dataset
5.5k/27,605English
Text-To-Music (TTM)2025SONICS
Paper
48,090/49,074English
Text-To-Music (TTM)2025SingNet
Paper Dataset
2,963.4 hoursMulti-lingual
Codec-based Speech
Generation (CoSG)
2024CodecFake
Paper Dataset
42,752/45,045English
Codec-based Speech
Generation (CoSG)
2024ALM-ADD
Paper Dataset
123/210English
Codec-based Speech
Generation (CoSG)
2024Codecfake
Paper Dataset
132,277/925,939English, Chinese
Codec-based Speech
Generation (CoSG)
2025ST-Codecfake
Paper Dataset
13,228/145,778English, Chinese
Codec-based Speech
Generation (CoSG)
2025CodecFake+
Paper
90,163/1,423,894English

Audio Preprocessing

Commonly Used Noise Datasets

⬆ Back to top

DatasetDescription
MUSAN
Dataset
A corpus of music, speech and noise
RIR
Dataset
A database of simulated and real room impulse responses, isotropic and point-source noises. The audio files in this data are all in 16k sampling rate and 16-bit precision.
NOIZEUS
Dataset
Contains 30 IEEE sentences (generated by three male and three female speakers) corrupted by eight different real-world noises at different SNRs. Noises include suburban train noise, murmur, car, exhibition hall, restaurant, street, airport and train station noise.
NoiseX-92
Dataset
All noises are obtained with a duration of 235 seconds, a sampling rate of 19.98 KHz, an analogue-to-digital converter (A/D) with 16 bits, an anti-alias filter and no pre-emphasis stage. Fifteen noise types are included.
DEMAND
Dataset
Multi-channel acoustic noise database for diverse environments.
ESC-50
Dataset
A tagged collection of 2000 environmental audios obtained from clips in Freesound.org, suitable for environmental sound classification. The dataset consists of 5-second-long recordings organised into 5 broad categories, each with 10 subcategories (40 examples per subcategory).
ESC
Dataset
Including the ESC-50, ESC-10, and ESC-US.
FSD50K
Dataset
An open dataset of human tagged sound events containing 51,197 Freesound clips totalling 108.3 hours of multi-labeled audio, unequally distributed across 200 classes from the AudioSet Ontology.

Audio Enhancement Methods

⬆ Back to top

MethodDescription
SpecAugment
Paper Code
Enhancement strategies include time warping, frequency masking and time masking
WavAugment
Paper Code
Enhancement strategies include pitch randomization, reverberation, additive noise, time dropout (temporal masking), band reject and clipping
RawBoost
Paper Code
Enhancement strategies include linear and non-linear convolutive noise, impulsive signal-dependent additive noise and stationary signal-independent additive noise

Feature Extraction

Handcrafted Feature-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork FrameworkLoss FunctionEER (%)t-DCF
Detecting spoofing attacks using VGG and SincNet: BUT-Omilia submission to ASVspoof 2019 challenge
Paper Code
—CQT, Power SpectrumVGG, SincNetCELA: 8.01 (4)
PA: 1.51 (2)
LA: 0.208 (4)
PA: 0.037 (1)
Long-term high frequency features for synthetic speech detection
Paper
Cafe, White and Street NoiseICQC, ICQCC, ICBC, ICLBCDNNCELA: 7.78 (3)LA: 0.187 (3)
Voice spoofing countermeasure for logical access attacks detection
Paper
—ELTP-LFCCDBiLSTM—LA: 0.74 (1)LA: 0.008 (1)
Voice spoofing detector: A unified anti-spoofing framework
Paper
—ATP-GTCCSVMHamming
Distance
LA: 0.75 (2)
PA: 1.00 (1)
LA: 0.050 (2)
PA: 0.064 (2)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA or PA scenario, and bolded values are the best results for that scenario.

Hybrid Feature-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
Light convolutional neural network with feature genuinization for detection of synthetic speech attacks
Paper
—CQT-based LPSLCNN—LA: 4.07 (11)LA: 0.102 (10)
Siamese convolutional neural network using gaussian probability feature for spoofing speech detection
Paper
—LFCCSiamese CNNCELA: 3.79 (10)
PA: 7.98 (5)
LA: 0.093 (5)
PA: 0.195 (2)
Generalization of audio deepfake detection
Paper
RIR and MUSANLFBResNet18LCMLLA: 1.81 (4)LA: 0.052 (4)
Continual learning for fake audio detection
Paper
—LFCCLCNN, DFWFSimilarity LossLA: 7.74 (15)
PA: 8.85 (6)
—
Partially-connected differentiable architecture search for deepfake and spoofing detection
Paper Code
Frequency MaskLFCCPC-DARTSWCELA: 4.96 (12)LA: 0.091 (8)
One-class learning towards synthetic voice spoofing detection
Paper Code
—LFCCResNet18OC-SoftmaxLA: 2.19 (7)LA: 0.059 (5)
Replay and synthetic speech detection with res2net architecture
Paper Code
—CQTSE-Res2Net50BCELA: 2.50 (8)
PA: 0.46 (2)
LA: 0.074 (7)
PA: 0.012 (2)
An empirical study on channel effects for synthetic voice spoofing countermeasure systems
Paper Code
Telephone Codecs, and Device/Room Impulse Responses (IRs).LFCCLCNN, ResNet-OCOC-Softmax, CELA: 3.92 (10)—
Efficient attention branch network with combined loss function for automatic speaker verification spoof detection
Paper Code
SpecAug, Attention MaskLFCCEfficientNet-A0, SE-Res2Net50WCE, Triplet LossLA: 1.89 (6)
PA: 0.86 (4)
LA: 0.507 (11)
PA: 0.024 (4)
Resmax: Detecting voice spoofing attacks with residual network and max feature map
Paper
—CQTResMaxBCELA: 2.19 (7)
PA: 0.37 (1)
LA: 0.060 (6)
PA: 0.009 (1)
Synthetic voice detection and audio splicing detection using se-res2net-conformer architecture
Paper
Adding noise according to a signal-to-noise ratio of 15dB or 25dBCQTSE-Res2Net34-ConfromerCELA: 1.85 (5)LA: 0.060 (6)
Fastaudio: A learnable audio front-end for spoof speech detection
Paper Code
—L-VQTL-DenseNetNLLLossLA: 1.54 (3)LA: 0.045 (3)
Learning from yourself: A self-distillation method for fake speech detection
Paper
—LPS, F0ECANet, SENetA-SoftmaxLA: 1.00 (2)
PA: 0.65 (3)
LA: 0.031 (2)
PA: 0.017 (3)
How to boost anti-spoofing with x-vectors
Paper
—LFCC, MFCCTDNN, SENet34LCMLLA: 0.83 (1)LA: 0.024 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA or PA scenario, and bolded values are the best results for that scenario.

End-to-end Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
A light convolutional GRU-RNN deep feature extractor for asv spoofing detection
Paper
—LC-GRNNPLDA—LA: 6.28 (13)
PA: 2.23
LA: 0.152 (10)
PA: 0.061
Rw-resnet: A novel speech anti-spoofing model using raw waveform
Paper
—1D Convolution Residual BlockResNetCELA: 2.98 (11)LA: 0.082 (9)
Raw differentiable architecture search for speech deepfake and spoofing detection
Paper Code
Masking FilterSinc FilterPC-DARTSP2SGradLA: 1.77 (10)LA: 0.052 (7)
Towards end-to-end synthetic speech detection
Paper Code
—DNNRes-TSSDNet, Inc-TSSDNetWCELA: 1.64 (9)LA: 0.048 (6)
End-to-end anti-spoofing with RawNet2
Paper Code
—Sinc FilterRawNet2CELA: 1.12 (5)LA: 0.033 (3)
Long-term variable Q transform: A novel time-frequency transform algorithm for synthetic speech detection
Paper
—FastAudio filterX-vector, ECAPA-TDNN—LA: 1.54 (7)LA: 0.045 (5)
Fully automated end-to-end fake audio detection
Paper
Sinc FilterWav2Vec2light-DARTSComparative lossLA: 1.08 (4)—
Audio anti-spoofing using a simple attention module and joint optimization based on additive angular margin loss and meta-learning
Paper
—Sinc FilterRawNet2, SimAMAAM Softmax, MSELA: 0.99 (3)LA: 0.029 (2)
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks
Paper Code
—Sinc FilterRawNet2, MGO, HS-GALCELA: 0.83 (2)LA: 0.028 (1)
Ai-synthesized voice detection using neural vocoder artifacts
Paper Code
Resampling, Noise AdditionSinc FilterRawNet2CE, SoftmaxLA: 4.54 (12)—
To-RawNet: Improving rawnet with tcn and orthogonal regularization for fake audio detection
Paper
RawBoostSinc FilterRawNet2, TCNCE, Orthogonal LossLA: 1.58 (8)—
Speaker-Aware Anti-spoofing
Paper
—Sinc FilterAASIST, M2S ConverterCELA: 1.13 (6)LA: 0.038 (4)
Spoofing attacker also benefits from self-supervised pretrained model
Paper
—HuBERT, WavLMResidual block, Conv-TasNetAAM softmaxLA: 0.44 (1)—
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Feature Fusion-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Feature ExtractionNetwork StructureLoss FunctionEER (%)
Voice spoofing countermeasure for synthetic speech detection
Paper
GTCC, MFCC, Spectral Flux, Spectral CentroidBi-LSTM—LA: 3.05 (4)
Combining automatic speaker verification and prosody analysis for synthetic speech detection
Paper
MFCC, Mel-SpectrogramECAPA-TDNN, Prosody EncoderBCELA: 5.39 (5)
Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
Paper
Sinc Filter, Wav2Vec2AASISTContrastive Loss, WCE—
Overlapped frequency-distributed network: Frequency-aware voice spoofing countermeasure
Paper
Mel-Spectrogram, CQTLCNN, ResNet—LA: 1.35 (2)
PA: 0.35
Detection of cross-dataset fake audio based on prosodic and pronunciation features
Paper
Phoneme Feature, Prosody Feature, Wav2Vec2LCNN, Bi-LSTMCTCLA: 1.58 (3)
Betray oneself: A novel audio deepfake detection model via mono-to-stereo conversion
Paper Code
Sinc FilterAASIST, M2S ConverterCELA: 1.34 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Network Training

Multi-task Learning-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Feature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
Multi-task learning in utterance-level and segmental-level spoof detection
Paper
LFCCSELCNN, Bi-LSTMP2SGrad——
SA-SASV: An end-to-end spoof-aggregated spoofing-aware speaker verification system
Paper Code
Fbanks, Sinc FilterECAPA-TDNN, ARawNetBCE, AAM Softmax, CELA: 4.86 (4)—
STATNet: Spectral and temporal features based multi-task network for audio spoofing detection
Paper
Sinc FilterRawNet2, TCM, SCMCELA: 2.45 (3)LA: 0.062 (2)
A probabilistic fusion framework for spoofing aware speaker verification
Paper Code
Mel Filter, Sinc FilterECAPA-TDNN, AASISTBCELA: 1.53 (2)—
DSVAE: Interpretable disentangled representation for synthetic speech detection
Paper
SpectrogramVAEKL Divergence Loss, BCELA: 6.56 (5)—
End-to-end dual-branch network towards synthetic speech detection
Paper Code
LFCC, CQTDual-Branch NetworkClassification Loss, Fake Type Classification LossLA: 0.80 (1)LA: 0.021 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Reference

⬆ Back to top

If you find our repository useful for your research, please cite it as follows:

@article{xu2024,
  title={Research progress on speech deepfake and its detection techniques},
  author={Xu Yuxiong, Li Bin, Tan Shunquan, Huang Jiwu},
  journal={Journal of Image and Graphics},
  volume={29},
  number={8},
  pages={2236--2268},
  year={2024}
}

Statement

⬆ Back to top

The purpose of this project is to establish a database based on audio deepfake detection, solely for the purpose of communication and learning. All the content collected in this project is sourced from journals and the internet, and we express sincere gratitude to the researchers and authors who have published related research achievements. In the event of a complaint of copyright infringement, the content will be removed as appropriate.

Significant stargazers

Xun Liu

53 followers · starred Sep 2024

Garima

14 followers · starred May 2025

media-sec-lab/Audio-Deepfake-Detection

Research progress on speech deepfake detection: Relevant datasets aggregated from the review literature and publicly available codes

324

60 commits

updated Jun 4, 2025

See the code

README

[!IMPORTANT]

A curated collection of papers and resources on Audio Deepfake Detection (ADD).

Please refer to our survey "Research progress on speech deepfake and its detection techniques" for the detailed contents. Paper page

Please let us know if you discover any mistakes or have suggestions by emailing us: xuyuxiong2022@email.szu.edu.cn

Table of contents

What's New

Survey

⬆ Back to top

  • 【2022-04】-【ADD-Survey】-【Ben Gurion University】
    • A study on data augmentation in voice anti-spoofing
    • Author(s): Ariel Cohen, Inbal Rimon, Eran Aflalo, Haim H. Permuter
    • Paper Code
  • 【2022-05】-【ADD-Survey】-【King Saud University】
    • A review of modern audio deepfake detection methods: challenges and future directions
    • Author(s): Zaynab Almutairi, and Hebah Elgibreen
    • Paper
  • 【2023-01】-【ADD-Survey】-【University of Maryland Baltimore County】
    • Audio deepfakes: A survey
    • Author(s): Zahra Khanjani, Gabrielle Watson and Vandana P. Janeja
    • Paper
  • 【2023-02】-【ADD-Survey】-【Oakland University】
    • Battling voice spoofing: a review, comparative analysis, and generalizability evaluation of state-of-the-art voice spoofing counter measures
    • Author(s): Awais Khan, Khalid Mahmood Malik, James Ryan, and Mikul Saravanan
    • Paper Code
  • 【2023-04】-【ADD-Survey】-【Panjab University】
    • Review of audio deepfake detection techniques: Issues and prospects
    • Author(s): Abhishek Dixit, Nirmal Kaur, Staffy Kingra
    • Paper
  • 【2023-07】-【ADD-Survey】-【Indian Institute of Technology】
    • Uncovering the deceptions: an analysis on audio spoofing detection and future prospects
    • Author(s): Rishabh Ranjan, Mayank Vatsa, Richa Singh
    • Paper
  • 【2023-08】-【ADD-Survey】-【CASIA】
    • Audio deepfake detection: a survey
    • Author(s): Jiangyan Yi, Chenglong Wang, Jianhua Tao, Xiaohui Zhang, Chu Yuan Zhang, and Yan Zhao
    • Paper
  • 【2024-02】-【Multimodal-Survey】-【Purdue University, Nanchang University】
    • Detecting Multimedia Generated by Large AI Models: A Survey
    • Author(s): Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu
    • Paper Code
  • 【2024-04】-【ADD-Survey】-【Toronto Metropolitan University】
    • A Survey on Speech Deepfake Detection
    • Author(s): Menglu Li, Yasaman Ahmadiadli, Xiao-Ping Zhang
    • Paper
  • 【2024-05】-【Multimodal-Survey】-【Cool Large Language Models Research Group, Renmin University of China】
    • Fake Artificial Intelligence Generated Contents (FAIGC): A Survey of Theories, Detection Methods, and Opportunities
    • Author(s): Xiaomin Yua, Yezhaohui Wanga, Yanfang Chenb, Zhen Tao, Dinghao Xi, Shichao Song, and Simin Niu
    • Paper
  • 【2024-09】-【ADD-Survey】-【Austrian Institute of Technology】
    • A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
    • Author(s): Lam Pham, Phat Lam, Tin Nguyen, Hieu Tang, Huyen Nguyen, Alexander Schindler, Hai Canh Vu
    • Paper Code
  • 【2024-11】-【Multimodal-Survey】-【University of Bucharest】
    • Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
    • Author(s): Forinel-Alin Croitoru, Andrei-Iulian Hıˆji, Vlad Hondru, Nicolae Cat ̆alin Ristea, Paul Irofti, Marius Popescu, Cristian Rusu, Radu Tudor Ionescu, Fahad Shahbaz Khan, and Mubarak Shah
    • Paper Code
  • 【2024-11】-【Multimodal-Survey】-【University College Dublin】
    • Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey
    • Author(s): Hong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac
    • Paper
  • 【2024-12】-【ADD-Survey】-【Imperial College London, Technical University of Munich】
    • From Audio Deepfake Detection to AI-Generated Music Detection——A Pathway and Overview
    • Author(s): Yupei Li, Manuel Milling, Lucia Specia, Björn W. Schuller
    • Paper
  • 【2025-02】-【Multimodal-Survey】-【BUPT, University of California, CASIA】
    • Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
    • Author(s): Yueying Zou, Peipei Li, Zekun Li, Huaibo Huang, Xing Cui, Xuannan Liu, Chenghanyu Zhang, Ran He
    • Paper

Top Repositories

⬆ Back to top

  • Audio Large Language Models
    • Resources on Audio Large Language Models, including datasets, methods, benchmarks, and studies.
    • Repo
  • awesome-fake-audio-detection
    • A list of tools, papers and code related to Fake Audio Detection.
    • Repo
  • ASVspoof Challenge Official Repository

Audio Large Model

⬆ Back to top

ModelPublisherYearsAchievable Tasks
AudioLM
Paper Website Code
Google2022.091. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation space.
2. Speech continuation, Acoustic generation, Unconditional generation, Generation without semantic tokens, and Piano continuation.
VALL-E
Paper Website
Microsoft2023.011. Simply record a 3-second registration of an unseen speaker to create a high-quality personalised speech.
2. VALL-E X: Cross-lingual speech synthesis.
USM
Website
Google2023.031. ASR beyond 100 languages.
2. Downstream ASR tasks.
3. Automated Speech Translation (AST).
SpeechGPT
Website
Fudan University2023.051. Perceive and generate multi-modal contents.
2. Spoken dialogue LLM with strong human instruction.
Pengi
Paper Website
Microsoft2023.051. an Audio Language Model that leverages Transfer Learning by framing all audio tasks as text-generation tasks.
2. The unified architecture of Pengi enables open-ended tasks and close-ended tasks without any additional fine-tuning or task-specific extensions.
VoiceBox
Website
Meta2023.061. Synthesize speech across six languages.
2. Remove transient noise.
3. Edit content.
4. Transfer audio style within and across languages.
5. Generate diverse speech samples.
AudioPaLM
Paper Website
Google2023.061. Speech-to-speech translation.
2. Automatic Speech Recognition (ASR).

Datasets

⬆ Back to top

Attack TypesYearsDatasetNumber of Audio
(Subdataset:Real/Fake)
Language
TTS2019FoR
Paper Dataset
111,000/87,000English
TTS2021WaveFake
Paper Dataset
16,283/117,985English, Japanese
TTS2021Half-truth
Paper
53,612/107,224Chinese
TTS2021HAD
Paper
53,612/107,224Chinese
TTS2021FakeAVCeleb
Paper Dataset
10,209/11,335English
TTS2022ADD 2022
Paper
LF: 5,619/46,067
PF: 5,319/46,419
FG-D: 5,319/46,206
Chinese
TTS2022CMFD
Paper Dataset
Chinese: 1,800/1,000
English: 1,800/1,000
English, Chinese
TTS2022In-the-Wild
Paper Dataset
19,963/11,816English
TTS2022CFAD
Paper Dataset
38,600/77,200Chinese
TTS2022Psynd
Paper Dataset
2,294English
TTS2022TIMIT-TTS
Paper Dataset
0/79,120English
TTS2023ODSS
Paper Dataset
11,032/18,993English, German,
and Spanish
TTS2024MLAAD
Paper Dataset
-/76,000Multi-lingual
TTS2024CD-ADD
Paper
300 hoursEnglish
TTS2024DiffSSD
Paper Dataset
24,226/70,000English
TTS2024ACCENT
Paper
53,651/192,461Multi-lingual
TTS2024SpoofCeleb
Paper Dataset
250k+English
TTS2024LlamaPartialSpoof
Paper Dataset
10,573/33,479English
TTS2024FakeSound
Paper Dataset
-/3,798English
TTS2024DFADD
Paper Dataset
44,455/163,500English
TTS2024SONAR
Paper Dataset
-/2,274Multi-lingual
TTS20256KSFx
Paper Dataset
-/6,000-
TTS2025BangalFake
Paper Dataset
12,260/13,260Bengali
Replay2017ASVspoof 2017
Paper Dataset
3,565/14,465English
Replay2019ReMASC
Paper Dataset
9,240/45,472English, Chinese,
Hindi
Replay2019VSDC
Paper Dataset
1,687/11,772English
Replay2024POLIPHONE
Paper Dataset
41,941English
TTS
and VC
2015AVspoof
Paper Dataset
LA: 15,504/120,480
PA: 15,504/14,465
English
TTS
and VC
2015ASVspoof 2015
Paper Dataset
16,651/246,500English
TTS
and VC
2021FMFCC-A
Paper Dataset
10,000/40,000Chinese
TTS
and VC
2022SceneFake
Paper Dataset
19,838/64,642English
TTS
and VC
2022EmoFake
Paper
35,000/53,200English, Chinese
TTS
and VC
2023PartialSpoof
Paper Dataset
12,483/108,978English
TTS
and VC
2023ADD 2023
Paper
FG-D: 172,819/113,042
RL: 55,468/65,449
AR: 14,907/95,383
Chinese
TTS
and VC
2023DECRO
Paper Dataset
Chinese: 21,218/41,880
English: 12,484/42,799
English, Chinese
TTS
and VC
2023HABLA
Paper Dataset
22,000/58,000Spanish
TTS
and VC
2024DeepFakeVox-HQ
Paper
693k/643kEnglish
TTS
and VC
2024VoiceWukong
Paper Dataset
5,300/413,400English, Chinese
TTS
and VC
2024Speech-Forensics
Paper Dataset
13,100/7,452English
TTS
and VC
2024VoiceEdit
Paper
-Multi-lingual
TTS
and VC
2024RFP
Paper Dataset
28,115/74,199English
TTS
and VC
2025MADD
Paper
60,000/129,990Multi-lingual
TTS
and VC
2025XMAD-Bench
Paper Dataset
414,858Multi-lingual
TTS
and Vocoder
2024Diffuse or Confuse
Paper Dataset
131,000/183,400English
TTS
and Vocoder
2025ShiftySpeech
Paper Dataset
3,000+ hoursEnglish, Chinese,
Japanese
TTS, VC
and Replay
2019ASVspoof 2019
Paper Dataset
LA: 12,483/108,978
PA: 28,890/189,540
English
TTS, VC
and Replay
2021ASVspoof 2021
Paper Dataset
LA: 18,452/163,114
PA: 126,630/816,480
PF: 14,869/519,059
English
TTS, VC
and Vocoder
2024SpeechFake
Paper
3,000+ hoursEnglish
VC, Replay
and Adversarial
2024VSASV
Paper Dataset
164,000/174,000Multi-lingual
TTS, VC
and Adversarial
2024ASVspoof 5
Paper Dataset
188,819/ 815,262English
Voice Cloning2021RTVCSpoof
Paper
3,284/4,843English
Voice Cloning2024Kratika Dataset
Paper
24,226/25,000English
Vocoder2022Yan Dataset
Paper
8,200/63,200Chinese
Vocoder2023LibriSeVoc
Paper Dataset
13,201/79,206English
Vocoder2023Voc.v1-v4
Paper Dataset
2,580/10,320English
Vocoder2024MLADDC
Paper Dataset
80k/160kMulti-lingual
Vocoder2024CVoiceFake
Paper Dataset
23,544/91,700Multi-lingual
Impersonation2024IPAD
Paper
5,170/18,874Chinese
Text-To-Music (TTM)2024FSD
Paper Dataset
200/500Chinese
Text-To-Music (TTM)2024SingFake
Paper Dataset
634/671Multi-lingual
Text-To-Music (TTM)2024CtrSVDD
Paper Dataset
32,312/188,486Multi-lingual
Text-To-Music (TTM)2024FakeMusicCaps
Paper Dataset
5.5k/27,605English
Text-To-Music (TTM)2025SONICS
Paper
48,090/49,074English
Text-To-Music (TTM)2025SingNet
Paper Dataset
2,963.4 hoursMulti-lingual
Codec-based Speech
Generation (CoSG)
2024CodecFake
Paper Dataset
42,752/45,045English
Codec-based Speech
Generation (CoSG)
2024ALM-ADD
Paper Dataset
123/210English
Codec-based Speech
Generation (CoSG)
2024Codecfake
Paper Dataset
132,277/925,939English, Chinese
Codec-based Speech
Generation (CoSG)
2025ST-Codecfake
Paper Dataset
13,228/145,778English, Chinese
Codec-based Speech
Generation (CoSG)
2025CodecFake+
Paper
90,163/1,423,894English

Audio Preprocessing

Commonly Used Noise Datasets

⬆ Back to top

DatasetDescription
MUSAN
Dataset
A corpus of music, speech and noise
RIR
Dataset
A database of simulated and real room impulse responses, isotropic and point-source noises. The audio files in this data are all in 16k sampling rate and 16-bit precision.
NOIZEUS
Dataset
Contains 30 IEEE sentences (generated by three male and three female speakers) corrupted by eight different real-world noises at different SNRs. Noises include suburban train noise, murmur, car, exhibition hall, restaurant, street, airport and train station noise.
NoiseX-92
Dataset
All noises are obtained with a duration of 235 seconds, a sampling rate of 19.98 KHz, an analogue-to-digital converter (A/D) with 16 bits, an anti-alias filter and no pre-emphasis stage. Fifteen noise types are included.
DEMAND
Dataset
Multi-channel acoustic noise database for diverse environments.
ESC-50
Dataset
A tagged collection of 2000 environmental audios obtained from clips in Freesound.org, suitable for environmental sound classification. The dataset consists of 5-second-long recordings organised into 5 broad categories, each with 10 subcategories (40 examples per subcategory).
ESC
Dataset
Including the ESC-50, ESC-10, and ESC-US.
FSD50K
Dataset
An open dataset of human tagged sound events containing 51,197 Freesound clips totalling 108.3 hours of multi-labeled audio, unequally distributed across 200 classes from the AudioSet Ontology.

Audio Enhancement Methods

⬆ Back to top

MethodDescription
SpecAugment
Paper Code
Enhancement strategies include time warping, frequency masking and time masking
WavAugment
Paper Code
Enhancement strategies include pitch randomization, reverberation, additive noise, time dropout (temporal masking), band reject and clipping
RawBoost
Paper Code
Enhancement strategies include linear and non-linear convolutive noise, impulsive signal-dependent additive noise and stationary signal-independent additive noise

Feature Extraction

Handcrafted Feature-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork FrameworkLoss FunctionEER (%)t-DCF
Detecting spoofing attacks using VGG and SincNet: BUT-Omilia submission to ASVspoof 2019 challenge
Paper Code
—CQT, Power SpectrumVGG, SincNetCELA: 8.01 (4)
PA: 1.51 (2)
LA: 0.208 (4)
PA: 0.037 (1)
Long-term high frequency features for synthetic speech detection
Paper
Cafe, White and Street NoiseICQC, ICQCC, ICBC, ICLBCDNNCELA: 7.78 (3)LA: 0.187 (3)
Voice spoofing countermeasure for logical access attacks detection
Paper
—ELTP-LFCCDBiLSTM—LA: 0.74 (1)LA: 0.008 (1)
Voice spoofing detector: A unified anti-spoofing framework
Paper
—ATP-GTCCSVMHamming
Distance
LA: 0.75 (2)
PA: 1.00 (1)
LA: 0.050 (2)
PA: 0.064 (2)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA or PA scenario, and bolded values are the best results for that scenario.

Hybrid Feature-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
Light convolutional neural network with feature genuinization for detection of synthetic speech attacks
Paper
—CQT-based LPSLCNN—LA: 4.07 (11)LA: 0.102 (10)
Siamese convolutional neural network using gaussian probability feature for spoofing speech detection
Paper
—LFCCSiamese CNNCELA: 3.79 (10)
PA: 7.98 (5)
LA: 0.093 (5)
PA: 0.195 (2)
Generalization of audio deepfake detection
Paper
RIR and MUSANLFBResNet18LCMLLA: 1.81 (4)LA: 0.052 (4)
Continual learning for fake audio detection
Paper
—LFCCLCNN, DFWFSimilarity LossLA: 7.74 (15)
PA: 8.85 (6)
—
Partially-connected differentiable architecture search for deepfake and spoofing detection
Paper Code
Frequency MaskLFCCPC-DARTSWCELA: 4.96 (12)LA: 0.091 (8)
One-class learning towards synthetic voice spoofing detection
Paper Code
—LFCCResNet18OC-SoftmaxLA: 2.19 (7)LA: 0.059 (5)
Replay and synthetic speech detection with res2net architecture
Paper Code
—CQTSE-Res2Net50BCELA: 2.50 (8)
PA: 0.46 (2)
LA: 0.074 (7)
PA: 0.012 (2)
An empirical study on channel effects for synthetic voice spoofing countermeasure systems
Paper Code
Telephone Codecs, and Device/Room Impulse Responses (IRs).LFCCLCNN, ResNet-OCOC-Softmax, CELA: 3.92 (10)—
Efficient attention branch network with combined loss function for automatic speaker verification spoof detection
Paper Code
SpecAug, Attention MaskLFCCEfficientNet-A0, SE-Res2Net50WCE, Triplet LossLA: 1.89 (6)
PA: 0.86 (4)
LA: 0.507 (11)
PA: 0.024 (4)
Resmax: Detecting voice spoofing attacks with residual network and max feature map
Paper
—CQTResMaxBCELA: 2.19 (7)
PA: 0.37 (1)
LA: 0.060 (6)
PA: 0.009 (1)
Synthetic voice detection and audio splicing detection using se-res2net-conformer architecture
Paper
Adding noise according to a signal-to-noise ratio of 15dB or 25dBCQTSE-Res2Net34-ConfromerCELA: 1.85 (5)LA: 0.060 (6)
Fastaudio: A learnable audio front-end for spoof speech detection
Paper Code
—L-VQTL-DenseNetNLLLossLA: 1.54 (3)LA: 0.045 (3)
Learning from yourself: A self-distillation method for fake speech detection
Paper
—LPS, F0ECANet, SENetA-SoftmaxLA: 1.00 (2)
PA: 0.65 (3)
LA: 0.031 (2)
PA: 0.017 (3)
How to boost anti-spoofing with x-vectors
Paper
—LFCC, MFCCTDNN, SENet34LCMLLA: 0.83 (1)LA: 0.024 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA or PA scenario, and bolded values are the best results for that scenario.

End-to-end Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Data AugmentationFeature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
A light convolutional GRU-RNN deep feature extractor for asv spoofing detection
Paper
—LC-GRNNPLDA—LA: 6.28 (13)
PA: 2.23
LA: 0.152 (10)
PA: 0.061
Rw-resnet: A novel speech anti-spoofing model using raw waveform
Paper
—1D Convolution Residual BlockResNetCELA: 2.98 (11)LA: 0.082 (9)
Raw differentiable architecture search for speech deepfake and spoofing detection
Paper Code
Masking FilterSinc FilterPC-DARTSP2SGradLA: 1.77 (10)LA: 0.052 (7)
Towards end-to-end synthetic speech detection
Paper Code
—DNNRes-TSSDNet, Inc-TSSDNetWCELA: 1.64 (9)LA: 0.048 (6)
End-to-end anti-spoofing with RawNet2
Paper Code
—Sinc FilterRawNet2CELA: 1.12 (5)LA: 0.033 (3)
Long-term variable Q transform: A novel time-frequency transform algorithm for synthetic speech detection
Paper
—FastAudio filterX-vector, ECAPA-TDNN—LA: 1.54 (7)LA: 0.045 (5)
Fully automated end-to-end fake audio detection
Paper
Sinc FilterWav2Vec2light-DARTSComparative lossLA: 1.08 (4)—
Audio anti-spoofing using a simple attention module and joint optimization based on additive angular margin loss and meta-learning
Paper
—Sinc FilterRawNet2, SimAMAAM Softmax, MSELA: 0.99 (3)LA: 0.029 (2)
AASIST: Audio anti-spoofing using integrated spectro-temporal graph attention networks
Paper Code
—Sinc FilterRawNet2, MGO, HS-GALCELA: 0.83 (2)LA: 0.028 (1)
Ai-synthesized voice detection using neural vocoder artifacts
Paper Code
Resampling, Noise AdditionSinc FilterRawNet2CE, SoftmaxLA: 4.54 (12)—
To-RawNet: Improving rawnet with tcn and orthogonal regularization for fake audio detection
Paper
RawBoostSinc FilterRawNet2, TCNCE, Orthogonal LossLA: 1.58 (8)—
Speaker-Aware Anti-spoofing
Paper
—Sinc FilterAASIST, M2S ConverterCELA: 1.13 (6)LA: 0.038 (4)
Spoofing attacker also benefits from self-supervised pretrained model
Paper
—HuBERT, WavLMResidual block, Conv-TasNetAAM softmaxLA: 0.44 (1)—
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Feature Fusion-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Feature ExtractionNetwork StructureLoss FunctionEER (%)
Voice spoofing countermeasure for synthetic speech detection
Paper
GTCC, MFCC, Spectral Flux, Spectral CentroidBi-LSTM—LA: 3.05 (4)
Combining automatic speaker verification and prosody analysis for synthetic speech detection
Paper
MFCC, Mel-SpectrogramECAPA-TDNN, Prosody EncoderBCELA: 5.39 (5)
Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation
Paper
Sinc Filter, Wav2Vec2AASISTContrastive Loss, WCE—
Overlapped frequency-distributed network: Frequency-aware voice spoofing countermeasure
Paper
Mel-Spectrogram, CQTLCNN, ResNet—LA: 1.35 (2)
PA: 0.35
Detection of cross-dataset fake audio based on prosodic and pronunciation features
Paper
Phoneme Feature, Prosody Feature, Wav2Vec2LCNN, Bi-LSTMCTCLA: 1.58 (3)
Betray oneself: A novel audio deepfake detection model via mono-to-stereo conversion
Paper Code
Sinc FilterAASIST, M2S ConverterCELA: 1.34 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Network Training

Multi-task Learning-based Forgery Detection

⬆ Back to top

PaperAudio Deepfake DetectionResults
Feature ExtractionNetwork StructureLoss FunctionEER (%)t-DCF
Multi-task learning in utterance-level and segmental-level spoof detection
Paper
LFCCSELCNN, Bi-LSTMP2SGrad——
SA-SASV: An end-to-end spoof-aggregated spoofing-aware speaker verification system
Paper Code
Fbanks, Sinc FilterECAPA-TDNN, ARawNetBCE, AAM Softmax, CELA: 4.86 (4)—
STATNet: Spectral and temporal features based multi-task network for audio spoofing detection
Paper
Sinc FilterRawNet2, TCM, SCMCELA: 2.45 (3)LA: 0.062 (2)
A probabilistic fusion framework for spoofing aware speaker verification
Paper Code
Mel Filter, Sinc FilterECAPA-TDNN, AASISTBCELA: 1.53 (2)—
DSVAE: Interpretable disentangled representation for synthetic speech detection
Paper
SpectrogramVAEKL Divergence Loss, BCELA: 6.56 (5)—
End-to-end dual-branch network towards synthetic speech detection
Paper Code
LFCC, CQTDual-Branch NetworkClassification Loss, Fake Type Classification LossLA: 0.80 (1)LA: 0.021 (1)
Note: "—" indicates not mentioned in the paper. Values in brackets in the experimental results are the ranking of each column in the LA scenario, and bolded values are the best results for that scenario.

Reference

⬆ Back to top

If you find our repository useful for your research, please cite it as follows:

@article{xu2024,
  title={Research progress on speech deepfake and its detection techniques},
  author={Xu Yuxiong, Li Bin, Tan Shunquan, Huang Jiwu},
  journal={Journal of Image and Graphics},
  volume={29},
  number={8},
  pages={2236--2268},
  year={2024}
}

Statement

⬆ Back to top

The purpose of this project is to establish a database based on audio deepfake detection, solely for the purpose of communication and learning. All the content collected in this project is sourced from journals and the internet, and we express sincere gratitude to the researchers and authors who have published related research achievements. In the event of a complaint of copyright infringement, the content will be removed as appropriate.

Significant stargazers

Xun Liu

53 followers · starred Sep 2024

Garima

14 followers · starred May 2025