TL;DR
Audio deepfakes (synthetic speech created by AI) have improved, but their effect on human trust in real speech was unknown. Researchers conducted the largest listening study to date, collecting 35,532 judgments from 1,768 participants evaluating 138 text-to-speech and voice conversion systems.
✦ Why It Matters
Engineers building voice authentication or speech verification systems must account for deepfake-induced skepticism when designing user interfaces and trust thresholds.
Key Takeaways
Full Summary
Recent advancements in audio deepfake technology have raised concerns about their impact on trust in real speech. A study involving 1,768 participants and 35,532 judgments assessed perceptions of 138 different text-to-speech and voice conversion systems.
Findings indicate that while participants' accuracy in identifying fake audio samples remained relatively unchanged (72.9% to 71.2%), their ability to recognize real audio samples dropped sharply from 72.7% to 64.1%. This suggests that the issue lies not in detecting fakes but in an increasing skepticism towards authentic speech.
Samples generated by commercial and autoregressive models were the hardest to identify, with detection rates between 61.3% and 65.9%, while traditional models were easier to spot (75.4% to 76.8%). An ML detector maintained over 94.5% accuracy across all conditions, highlighting the growing challenge of maintaining trust in genuine audio amidst rising deepfake technology.
Related