Phoneme-level framework for explainable speech deepfake detection
A phoneme-level framework using wav2vec 2.0 and HuBERT explains why speech is classified as real or fake. Other papers study robustness to real-world corruption and generalization to synthetic sound effects. A large-scale analysis of DETECT-3B-Omni confirms its independence from speech content and demographics.
How this story unfolded
3 days · 6 reports · from Jul 7
- Jul 7
DETECT-3B-Omni is Agnostic of Content and Demographicsarxiv.org
Doppelganger: Sound Effects and Their Synthetic Twinsarxiv.org
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluationarxiv.org
Measuring the Robustness of Audio Deepfake Detection under Real-World Corruptionarxiv.org
- Jul 8
- Jul 10
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Stable Diffusion user tests H3 model with Cheers-style script
- Reddit users share impressive image-to-video AI demos
- Reddit reminds users they can legally seed AI models via torrenting
- MiniMax H3 reverse-engineers paintings into basic forms
- OpenAI DevDay Exchange Seoul applications close Sept 4