Frontier AI firm Modulate has raised an additional $25 million funding from Future Ventures, with participation from Hyperplane and Lakestar, bringing its total funding to $60 million. Modulate builds audio-native AI models that enable machines to understand the nuances in human conversation, whether natural or fake.
It uses a proprietary Ensemble Listening Model (ELM) architecture to generate the individual AI models that power its Velma platform. Velma brings together more than 100 of these specialized models that understand and detect signals including emotion, tone, intent, synthetic speech and conversational behavior.
The firm currently analyzes more than 10 million hours of audio each month, with over 600 million hours analyzed in total. According to Modulate, its transcription and deepfake detection models both ranked #1 on public benchmarks like Hugging Face.
Modulate focuses on using AI to analyze voice rather than create voice. It claims its Velma platform results are twice as accurate as traditional LLMs and generate seven times fewer false positives. Since Velma operates in real time, applications built with it don’t simply understand what is happening, they can intervene while it is in process.
Typical applications are used to protect healthcare institutions from deepfake hackers; enhance voice AI agent’s emotion and empathy capabilities; reduce extremism and harassment on social platforms; observe and monitor voice agent performance; detect and stop child grooming voice conversations; and protect agents in high-risk scenarios from being identified through advanced voice masking.
“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript,” explains Carter Huffman, CEO and co-founder at Modulate. “We’re already using audio-native AI to protect organizations from deepfake attacks, help voice agents understand emotion and respond with more empathy, identify dangerous behavior in online conversations, and monitor whether voice agents are actually performing the way they’re supposed to.”
The new funding will be used to further develop Modulate’s team and infrastructure, and to expand the models, APIs, SDKs, integrations and deployment options available to the developers and partners that build voice applications; while continuing to invest in its own underlying research and technology.
Adversarial use of AI-assisted fraudulent deepfakes, and simple misuse of language for abusive purposes are now a fact of life. Detecting deepfakes and recognizing abuse in time to prevent harm is almost impossible without the use of AI. Building applications able to detect the misuse of voice is difficult. Modulate’s purpose is to assist in this.
“Developers shouldn’t have to rebuild the audio intelligence layer every time they create a new voice experience,” Huffman added. “Our mission is to build the models and infrastructure that let them focus on the application they want to create. The opportunity facing audio-native AI is expanding incredibly quickly. We’ve built the technology and proven it at scale, and this investment lets us grow the team and move faster to meet that demand.”
Related: EU to Crack Down on AI Deepfakes, Illicit Imagery and Hacking
Related: isVerified Emerges From Stealth With Voice Deepfake Detection Apps
Related: Creating Deepfakes Is Getting Easier Than Ever. Fighting Back May Take Even More AI
Related: The AI Arms Race: Deepfake Generation vs. Detection

