Voice analytics startup Modulate raises $25M for emotion, deepfake detection and policy suite

Boston company's platform runs more than 100 models for transcription, emotional analysis and audio authenticity checks

By LineZotpaper
Published
Read Time2 min
Sources2 outlets
Boston-based voice intelligence startup Modulate has raised $25 million in new funding for its platform, which uses an array of small models to offer enterprises transcription, emotional analysis, deepfake and AI music detection, and policy enforcement for voice agents in regulated industries.

The round was led by Future Ventures with participation from Hyperplane and Lakestar. Data from PitchBook indicated that the startup had raised $41 million in funding at a $170 million valuation prior to this round.

Modulate was founded in 2017 by Mike Pappas and Carter Huffman, who met as MIT physics undergrads. In its early days, the company focused on providing voice modulation for gaming, before developing a voice-based moderation tool. With the rise of voice AI models, the company is now concentrating on detecting different sorts of AI audio generation and analysing the intent behind a person's words.

"Our insight into the voice AI space is that a lot of folks are doing transcription, but there's not really any capability out there that gets the full nuance and full understanding of a conversation, which is so important when you're talking to another human being," Huffman said on a call with TechCrunch.

The company today runs more than 100 models, largely categorised into two sections. Signal extraction models are designed to understand vocal emotion, tone, language and synthetic audio.

The funding follows a popular trend among investors in the growing voice AI industry: backing companies that are trying to make AI voices sound more human. It also puts Modulate in competition with companies detecting the intent behind human conversation, and with those trying to protect people and companies from deepfake calls, as voice cloning has become easier.

§

Analysis

Why This Matters

  • Voice AI is expanding beyond transcription into interpretation, with investors backing tools that analyse emotion, tone and intent rather than just words.
  • Deepfake and AI-generated audio detection is a practical concern for enterprises and regulated industries facing cloned-voice fraud.
  • The round signals growing demand for policy enforcement as voice agents are deployed in sectors with compliance obligations.

Background

Modulate began in 2017 with voice modulation for gaming before pivoting to voice-based moderation. As AI-generated audio has become harder to distinguish from human speech, the company has shifted again, toward detecting synthetic audio and interpreting conversation. It now positions itself as serving regulated industries where voice agents must comply with policy, and where organisations need to know whether they are talking to a human.

Key Perspectives

Modulate: The company argues that transcription alone misses the full nuance and understanding of a conversation, which it says is essential when talking to another human being.

Investors: Future Ventures led the round with Hyperplane and Lakestar, reflecting a broader trend of backing startups that aim to make AI voices sound more human.

Competitors and skeptics: Rivals include companies analysing conversational intent and those protecting people from deepfake calls. Skeptics may question whether small-model analysis can keep pace with rapidly improving voice cloning and AI generation tools, and whether claims about reading emotion from audio hold up in practice.

What to Watch

  • How quickly Modulate converts its 100-plus model suite into enterprise contracts in regulated industries.
  • The competitive response from intent-analysis and deepfake-detection rivals.
  • Whether regulatory pressure on voice agents and AI-generated audio accelerates demand for policy enforcement tools.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.