The ability of artificial intelligence (AI) to generate synthetic content continues to grow. From images and videos to text and voice, AI-generated creations are increasingly difficult to distinguish from the real thing.
In particular, AI-generated speech has raised significant concerns due to its potential to be used for disinformation, fraud, and other malicious purposes.
Here we explore whether it is possible to detect AI-generated voices and how the challenges associated with this emerging technology are being addressed.
The rise of AI-generated voice
Speech synthesis has seen remarkable progress thanks to deep neural networks and other advances in machine learning. Tools like FakeYou and PlayHT have made it easy to create realistic synthetic voices, making it accessible to a wide range of users.
These systems can generate voices from text (TTS) or transform one person’s voice into another’s (voice conversion), retaining emotional nuances and breathing patterns of the original speech.
However, this technology has also been used to create audio “deepfakes“, where voices of public figures are falsified to spread disinformation or carry out scams.
One recent case involved a hoax call that appeared to be from US President Joe Biden, discouraging Democratic voters in the New Hampshire primary. This incident underscores the urgent need for reliable tools to detect AI-generated audio.
Challenges in detecting AI generated audio
Detecting AI-generated voices is inherently more complicated than detecting fake images or videos. Audio is a one-dimensional and ephemeral medium, making it difficult to review and analyze for signals of AI generation.
According to Manjeet Rege, director of the Center for Applied Artificial Intelligence at the University of St. Thomas, the audio lacks visual context and clues that can help identify fakes.
Current detection methods rely heavily on machine learning, where models are trained with large data sets of real and synthetic voices. That is, using AI to detect content generated by other AIs.
These models look for patterns and features that can distinguish AI-generated audio from real audio. However, these systems face several obstacles:
- Audio Variability: Audio can easily degrade due to compression, background noise, and other factors, making accurate detection difficult.
- Constant updating: With new AI voice generation models released every week, detection systems must be continually updated to identify subtle differences between real and generated voices.
- Language limitations: Most current models focus on English, which means they may not be effective at detecting audio generated in Spanish or other languages.
Is there anything effective to detect voice deepfakes?
Various institutions and companies are developing tools to detect AI-generated voices, although with mixed results. An experiment conducted by NPR (US public radio) tested 3 deepfake detection tools: Pindrop Security, AI or Not and AI Voice Detector.
Results varied significantly, with some tools failing to identify AI-generated clips or wrongly labeling real voices as fake.
Researchers from the University of Granada (UGR) have developed a pioneering system to discern whether an audio is real or generated by AI, integrating specific models for voices of personalities frequent in misinformation.
This tool, part of the RTVE-UGR Chair, aims to be an advanced solution for verifying news and combating disinformation.
Another study at the University at Buffalo tested 14 audio deepfake detection tools and found that none were completely reliable. Tools tested included DeepFake-o-meter and AI or Not, with results varying widely depending on audio and testing conditions.
Recommendations and combined approaches
Since current tools are not completely reliable, experts recommend a combined approach to detecting audio deepfakes. This includes the use of multiple detection methods and additional techniques.
For example, whenever possible, audio cross-checking should be performed. This involves confirming the authenticity of the audio by verifying it from independent and reliable sources.
Another strategy involves paying attention to subtle details, such as listening for breathing irregularities, pauses, and intonations that could indicate AI manipulation.
Other strategies involve psychology, such as questioning the urgency of a request. For example, audios that urgently request personal or financial information should always be treated with suspicion, as this sense of urgency serves to prevent victims from verifying the information.
Big platforms like Meta and TikTok are developing technologies to label AI-generated content, which could also include audio in the future.
The future of voice detection created by AI
As AI speech generation technology continues to evolve, it is essential that detection tools keep pace. Collaboration between speech synthesis technology developers and detection experts can increase the accuracy and effectiveness of these tools.
Additionally, implementing policies and regulations that require identification of AI-generated content can help mitigate the impact of disinformation and other malicious uses.
In summary, although AI-generated speech detection presents numerous challenges, continued advances in research and development of targeted tools promise to improve our ability to identify and mitigate associated risks.
It is critical that society remains vigilant and adopts combined approaches to protect against the emerging threats posed by this rapidly evolving technology.
This post is also available in: