
Modulate Raises $25 Million to Expand Audio-Native AI
Published by AINave Editorial • Reviewed by Ramit
Modulate has raised $25 million to bring its audio-native models to more developers, with plans for new software development kits and APIs. Its Velma platform analyzes raw conversation audio, looking beyond a transcript for signals such as tone, emotion, intent and possible deepfakes. Some analysis can run while a call is still happening.
That difference matters for voice-agent teams: a transcript captures what someone said, while audio can also carry signals the words alone miss. Modulate says its models can combine those signals to flag possible fraud or a caller losing patience with an agent. Those are described capabilities, not a guarantee that the system will correctly detect every event in a live deployment.
A platform built from specialized audio models
Under Velma, Modulate’s Ensemble Listening Model architecture selects from more than 100 specialized audio models for each task and blends their results. The company says this design is up to 1,000 times more efficient than sending the same audio to one large model. That comparison is Modulate’s claim; the report does not specify the measurement conditions or what workloads it covers.
The product has roots in voice-chat moderation for online games, where Modulate’s models still look for harassment and child grooming. The company also serves health care institutions screening for callers impersonating staff with deepfake voices, and customers using its models to assess voice-agent performance. Modulate says its models now process more than 10 million hours of audio a month, with lifetime processing recently exceeding 600 million hours.
Funding expands access, but results need context
Future Ventures led the round, joined by returning investors Hyperplane and Lakestar. Modulate says it will also use the funding for research, engineering and developer-relations hiring, industry-specific models, partner integrations and additional deployment options. The planned tools suggest the company is trying to make audio analysis a reusable layer for different voice products, rather than a capability each team must build from scratch.
Modulate reports that its Velma Deepfake Detect model achieved 98.9% accuracy on public benchmark data. It also says the model, launched in March, leads the Speech Deepfake Arena, while its transcription model topped the Open ASR Leaderboard in July. These results describe performance on the named benchmarks, not a universal rate for real calls. The batch transcription API costs three cents per hour; that figure does not establish the cost of other products or deployments.
The useful distinction is clear: Velma is positioned to analyze how a conversation sounds as well as what it says. Whether that broader view is worth adding will depend on a team’s task and on how the models perform in its own operating conditions.






















