The World's Leading Intelligence & Artificial Intelligence Journal

Home / AI & Models / Beyond the Transcript: Modulate’s $25M Bet on Audio-Native Intelligence
AI & Models • Sep 28, 2026 • 6 min read

Beyond the Transcript: Modulate’s $25M Bet on Audio-Native Intelligence

Modulate has secured $25 million to pivot enterprise security from text-based transcription to real-time, audio-native signal processing. By analyzing the raw physics of human speech, the platform aims to neutralize deepfake threats and emotional manipulation before they impact corporate bottom lines.

Ajinkya Pawar

By Ajinkya Pawar

Head of Search & AI Intelligence • The AI NEWS

Beyond the Transcript: Modulate’s $25M Bet on Audio-Native Intelligence
Beyond the Transcript: Modulate’s $25M Bet on Audio-Native Intelligence

Key Developments & Executive Briefing

Executive Briefing
01

Series Expansion

Funding $25M

Modulate secures $25M in new capital, bringing total funding to $60M.

02

Signal Processing

Tech Shift Audio-Native

Moving beyond NLP transcripts to raw audio waveform analysis.

03

Deepfake Defense

Security Real-Time

Proactive detection of synthetic voice fraud in enterprise environments.

Beyond the Transcript: Why Velma Bypasses Text-Based Moderation

Modern enterprise security has long relied on the transcription of voice calls into text to identify risk. However, this process strips away the critical metadata of human communication—the tone, the hesitation, and the emotional resonance that define intent. While traditional AI optimization often focuses on text-based efficiency, Modulate’s approach suggests that the future of enterprise safety lies in raw audio signal analysis.

Their flagship platform, Velma, processes audio as a high-fidelity data stream rather than a static document. By analyzing the physics of the voice, Velma can detect anomalies that text-based models would miss entirely, such as the subtle artifacts of a synthetic voice or the physiological markers of a stressed caller.

Metric | Text-Based Transcription Analysis | Audio-Native Signal Processing
:--- | :--- | :---
Latency | High (requires conversion) | Ultra-Low (real-time)
Emotional Nuance | Lost | Preserved
Fraud Detection Accuracy | Moderate | High
Computational Overhead | Low | Moderate/High

The Synthetic Arms Race: Detecting Deepfakes in Real-Time

As voice cloning technology becomes democratized, the threat of social engineering has reached a fever pitch. Corporate security teams are no longer just fighting phishing emails; they are fighting synthetic voices that sound indistinguishable from their own executives. As the industry grapples with the ethics of AI music and synthetic generation, Modulate’s ability to detect synthetic audio becomes a critical tool for corporate integrity.

"Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript," says CEO Carter Huffman. By acting as a real-time firewall, Modulate’s models provide a necessary layer of defense that validates the authenticity of the speaker before a transaction or sensitive data transfer occurs.

Physics in the Hallway: The MIT Roots of Modulate’s Architecture

Modulate’s technical DNA is rooted in the halls of MIT, where founders Mike Pappas and Carter Huffman first crossed paths. Their collaboration began not in a boardroom, but over a complex physics problem that required a fundamental understanding of signal behavior. This shared background in physics has allowed them to approach AI not as a black-box software problem, but as a signal processing challenge.

Key technical milestones include:

  • Initial development of physics-based audio analysis algorithms.
  • Deployment of the Velma platform for real-time enterprise moderation.
  • Expansion into multi-modal detection, including deepfake and synthetic music identification.
  • Reaching a $60M total funding milestone to scale infrastructure.

Capitalizing on the Voice Interface Gold Rush

The $25M funding round, led by Future Ventures with participation from Hyperplane and Lakestar, signals a massive shift in how capital is allocated toward AI infrastructure. Modulate’s recent raise highlights how capital is flowing into specialized AI infrastructure that secures the next generation of voice-driven enterprise workflows. Investors are clearly betting that the next wave of enterprise productivity will be voice-first, and that security must evolve to match that medium.

Workflow Timeline: The Evolution of Modulate

  • Phase 1 (Prototype): Early research into physics-based audio signal processing at MIT.
  • Phase 2 (Initial Funding): Establishing the core team and building the foundational voice models.
  • Phase 3 (Velma Launch): Moving from research to a commercial-grade, real-time enterprise platform.
  • Phase 4 (Scale): The $25M injection to expand the platform's reach into regulated industries.