Full-Time
Real-time voice skins for gaming
$180k - $200k/yr
Cambridge, MA, USA
Hybrid
Hybrid: core in-office days with flexible remote options.
See people who can refer or advise you
Modulate provides real-time voice skin technology for online gaming and virtual communication. It licenses its voice-skin systems to game developers and online platforms (B2B), so players can modify their voices to sound like characters or create new ones, integrating with existing voice chat. The product works by capturing the user's voice and applying real-time voice modulation skins, allowing seamless swapping of voice identities during gameplay. Modulate differentiates itself by targeting game developers and communities with an emphasis on inclusivity and engagement, reducing intimidation for players who are uncomfortable with their natural voices and encouraging participation. The company’s goal is to expand immersive, personalized interactions in online gaming by offering scalable licensing options and premium voice-skin variants through subscriptions or one-time fees.
Company Size
51-200
Company Stage
Series A
Total Funding
$66M
Headquarters
Somerville, Massachusetts
Founded
2017
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Hybrid Work Options
Scam.ai and Modulate have partnered to integrate synthetic voice detection into Scam.ai's deepfake detection platform. The collaboration allows organisations to identify AI-generated image, video, and audio threats through a single unified system. Scam.ai previously offered detection capabilities for images, videos, and digital documents. By adding Modulate's voice detection models, the platform now addresses multimodal deepfake attacks that span multiple communication channels. Modulate's technology supports real-time and prerecorded audio analysis, reporting 98.9% accuracy and currently holds first place on the Hugging Face Speech Deepfake Detection Leaderboard. Scam.ai's visual detection models report 98.2% accuracy. The integrated voice detection capability is expected to launch in early September. Applications include identity verification, fraud prevention, contact centre security, and content moderation.
Scam.ai and Modulate partner to deliver unified image, video and voice Deepfake Detection. August 4, 2026 10:22 AM Gift Article Integration brings Modulate's industry-leading synthetic voice detection to the Scam.ai platform, helping organizations identify multimodal deepfake threats through a single customer experience BOSTON, MA / ACCESS Newswire / August 4, 2026 / Modulate, the frontier conversational voice intelligence company, and Scam.ai, a provider of deepfake and synthetic media detection technology, today announced a partnership that brings Modulate's synthetic voice detection models directly into the Scam.ai platform. Scam.ai has established deepfake detection capabilities across images, videos, and digital documents. By integrating Modulate's specialized synthetic voice detection models, Scam.ai will enable customers to expose the three primary forms of synthetic media - image, video, and audio - through one unified platform and workflow. The integration addresses a growing challenge for organizations as deepfake attacks expand beyond a single format or communication channel. A fraudulent interaction may begin with a cloned voice over the phone, move to a fabricated image or document, and conclude with a manipulated video or identity-verification attempt. Defenses that examine only one component of the interaction risk missing the broader attack. "Scammers stopped limiting themselves to one channel a long time ago, but many detection systems are still organized around individual media formats," said Dr. Ben (Simiao) Ren, Co-founder and CEO of Scam.ai. "By integrating Modulate's industry-leading models into Scam.ai, we are giving our customers a practical way to add synthetic voice detection to the platform and workflows they already use to analyze visual content. This gives fraud and security teams a more complete view of the interactions they are evaluating without deploying another standalone tool." Through the partnership, Scam.ai will offer Modulate's synthetic voice detection models as part of its own product portfolio and customer experience. Customers will be able to analyze live or prerecorded audio alongside images and videos, receiving confidence scores and detection signals that indicate whether a voice is synthetic or AI-generated. "Deepfake attacks do not distinguish the boundaries between audio, images and video, and the technology used to stop them can't afford to either," said Carter Huffman, CTO and co-founder of Modulate. "Voice is increasingly part of coordinated multi-media deepfake scams. A convincing cloned voice can establish urgency and trust, while a fabricated video, image or document reinforces the deception. Scam.ai understands that organizations need to evaluate the entire interaction, and this partnership puts voice detection directly into the platform across workflows their customers already use." Deepfake Scams Escalate, Costing Billions The partnership arrives as scams become increasingly communication-based, more sophisticated, and move between multiple channels. According to research from Gallup and the Stop Scams Alliance, an estimated 15.1 million U.S. adults were personally scammed in 2025, resulting in at least $68 billion in losses. The study also found that phone calls, text messages and email were each involved in 45% of scams, with phone calls serving as the primary communication method more frequently than any other channel. Half of scams crossed two or more communication methods, reinforcing the need for detection systems that can examine multiple forms of content as part of a connected interaction. "People are being asked to determine whether a voice, image or video is authentic at the exact moment a scammer is trying to manipulate them," Huffman added. "That is an adversarial problem, and detection cannot depend on whether someone thinks a voice sounds suspicious. Organizations need automated systems that can analyze synthetic-media signals, explain why content was flagged, and help people make better decisions before money, access or sensitive information changes hands." Two Industry-Leading Technologies, One Unified Deepfake Defense Modulate's synthetic voice detection technology supports real-time streaming and prerecorded audio. The model reports 98.9% accuracy and a 1.1% equal error rate, and holds first place on the Hugging Face Speech Deepfake Detection Leaderboard as of August 4, 2026. It returns confidence scores and detailed detection signals through the Modulate API built for integration into enterprise platforms and applications. Scam.ai's platform currently provides real-time analysis of AI-generated and manipulated images and videos, with its Eva-v1 models reporting 98.2% visual detection accuracy against Scam.ai's internal benchmark. The platform is designed to return confidence scores and manipulation analysis through a single API and customer interface. Together, Modulate and Scam.ai will enable customers to: * Detect synthetic and manipulated content across image, video and voice through one unified platform and seamless user experience. * Reduce reliance on disconnected, medium-specific detection tools. * Add synthetic voice detection to existing fraud prevention, authentication and content-verification workflows. * Use confidence scores and detection signals to prioritize high-risk content for additional review. Potential applications include identity verification and digital onboarding, financial fraud and payment authorization, executive and employee impersonation, contact center security, social media and user-generated content moderation, insurance claims, digital evidence verification and enterprise cybersecurity investigations. The integrated voice detection capability is expected to be available through Scam.ai in early September. Customers can access additional information, request a demonstration or discuss availability at www.scam.ai. About Scam.ai Scam.ai, developed by Reality Inc., provides AI-powered technology for detecting deepfakes, synthetic media and other AI-enabled threats. Its platform helps organizations analyze images, videos, documents and digital content for signs of manipulation or artificial generation through a unified API and customer experience. Scam.ai supports applications across financial services, identity verification, media, insurance, contact centers, online platforms and enterprise security. About Modulate Modulate is a voice intelligence company building AI models and APIs designed to understand real-world conversational audio at scale. Its technology combines speech recognition, acoustic analysis, and conversational context to deliver reliable, explainable, and cost-effective voice intelligence for developers and enterprises. Media Contact
Modulate will showcase its voice-native AI architecture at Ai4 2026, taking place 4-6 August at The Venetian in Las Vegas. The Boston-based company's voice intelligence platform, Velma, is powered by models that rank first on the Hugging Face Open ASR Leaderboard, outperforming solutions from Microsoft, OpenAI, and NVIDIA. CTO and co-founder Carter Huffman will present on 5 August about conversational blind spots that transcript-only monitoring systems miss, such as emotion and frustration signals. Modulate's technology analyses over 500 million hours of real-world audio to help organisations identify customer frustration, compliance issues, and escalation risks in live conversations. The company's transcription APIs cost between $0.025 and $0.06 per hour, making them up to 10 times less expensive than several leading commercial services.
Modulate to showcase voice-native AI architecture at Ai4 2026. July 21, 2026 7:18 AM Ai4 attendees can experience Modulate's #1-ranked industry leading voice intelligence AI model and hear from CTO and co-founder Carter Huffman on the hidden costs of conversational blind spots BOSTON, MA / ACCESS Newswire / July 21, 2026 / Modulate, the frontier conversational voice intelligence company, will showcase its leading voice-native AI architecture at Ai4 2026, taking place August 4-6, 2026, at The Venetian in Las Vegas. At booth 940, Modulate will showcase Velma, a voice intelligence platform that provides real-time conversational understanding, powered by Modulate's homegrown models, which rank #1 on the Hugging Face Open ASR Leaderboard. Powered by Modulate's Ensemble Listening Model (ELM) architecture, the API helps organizations identify costly conversational blind spots - like emotion, emphasis, and frustration signals - that transcript-only monitoring systems frequently miss. Speaking alongside industry luminaries like Geoffrey Hinton and Andrew Ng, Modulate CTO and co-founder Carter Huffman will present a solo talk titled "Voice Agent Supervision: Finding the Hidden Cost of Conversational Blind Spots." During the session, Huffman will explain where traditional AI fails to create expected ROI, and why measuring more than just technical accuracy is necessary to facilitate successful customer interactions. At Ai4, Modulate will demonstrate how its Velma platform provides a voice-native supervision layer for live conversations. By preserving the conversational signals that are often lost when audio is converted to text, Velma enables organizations to identify customer frustration, escalation risk, compliance issues, and other conversational blind spots before they become operational or financial liabilities. This enables enterprises to identify where voice agents may be: * Failing to recognize customer frustration or confusion * Responding inappropriately to urgency, vulnerability, or emotional distress * Repeating information without moving the interaction toward resolution * Missing emphasis that changes the meaning or importance of a customer's request * Making inaccurate claims or commitments * Violating organizational policies or approved communication guidelines * Escalating tension through an unsuitable tone or response * Creating avoidable transfers, callbacks, abandonment, or churn Modulate helps organizations translate these interaction-level failures into operational and financial metrics. Carter Huffman Reveals Where AI Fails to Create ROI at Ai4 2026 Modulate CTO and co-founder Carter Huffman will present "Voice Agent Supervision: Finding the Hidden Cost of Conversational Blind Spots" on Wednesday, August 5th from 4:05 to 4:25pm PDT in Palazzo Ballroom D. During the session, Huffman will explain why some of the most expensive voice-agent failures do not appear as traditional software defects. Instead, they emerge when an agent misses tone, emotion, emphasis, hesitation, frustration, or other conversational signals that were never captured in the transcript. "Voice agents can appear successful on a traditional scorecard while quietly creating friction throughout the conversation," said Huffman. "An agent may produce a factually acceptable response and still miss that the customer is confused, losing patience, or emphasizing that the issue is urgent. If enterprises cannot measure those moments, they cannot understand the true performance or ROI of their voice-agent deployments." Attendees will learn how to: * Identify conversational blind spots in voice-agent deployments * Treat tone and emphasis as measurable, first-class signals * Detect and correct friction while a conversation is still underway * Distinguish technical accuracy from genuine conversational understanding * Translate agent-level errors into financial and operational metrics * Build a clearer and more defensible ROI case for enterprise voice AI Huffman is CTO and co-founder of Modulate, where he leads the development of the company's purpose-built Voice Intelligence Engine. A physicist and machine-learning expert trained at MIT, with previous experience at NASA's Jet Propulsion Laboratory, Huffman has spent more than a decade developing AI systems capable of understanding real-world conversational audio beyond the words spoken. Ranked #1 on the Hugging Face Open ASR Leaderboard At Ai4, Modulate will also highlight its recent recognition as the #1 model on Hugging Face's Open ASR Leaderboard, outperforming models from Microsoft, OpenAI, NVIDIA, IBM, ElevenLabs, AssemblyAI, BosonAI, and other commercial and open-source providers. Modulate ranked first among 88 models evaluated across standardized datasets covering multiple domains, accents, speakers, and recording conditions. The milestone demonstrates that Modulate's voice-native architecture can provide the accurate, low-latency, and cost-efficient transcription foundation required for production voice applications. Modulate's transcription APIs are priced between $0.025 and $0.06 per hour, making them up to 10 times less expensive than several other leading commercial transcription services. Modulate trains its models using more than 500 million hours of noisy, real-world audio. The company developed its technology in demanding, high-scale voice environments where conversations are live, emotionally complex, overlapping, and rarely studio-quality. That foundation now supports enterprise applications across voice-agent supervision, customer experience, fraud prevention, trust and safety, compliance, and conversational intelligence. Meet Modulate at Ai4 Attendees can visit Modulate at Booth 940 to experience live demonstrations of Velma and learn how voice-native conversational intelligence can help enterprises supervise AI agents, improve customer experiences, detect emerging risk, and identify the value being lost through conversational blind spots. To schedule a press briefing, contact Kristin Canders at [email protected]. Join Modulate for Happy Hour Modulate and Axonis will co-host a Happy Hour at Ai4 on Tuesday, August 4, from 6:00-8:00 p.m. at SUGARCANE, The Venetian in Las Vegas. Attendees are invited to join both teams for drinks and conversation. Register here: https://luma.com/Axonis-Modulate-Ai4 About Modulate Modulate is a voice intelligence company building AI models and APIs designed to understand real-world conversational audio at scale. Its technology combines speech recognition, acoustic analysis, and conversational context to deliver reliable, explainable, and cost-effective voice intelligence for developers and enterprises.
Modulate to showcase voice-native AI architecture at Ai4 2026. July 21, 2026 Boston, MA - July 21, 2026 - Modulate, the frontier conversational voice intelligence company, will showcase its leading voice-native AI architecture at Ai4 2026, taking place August 4-6, 2026, at The Venetian in Las Vegas. At booth 940, Modulate will showcase Velma, a voice intelligence platform that provides real-time conversational understanding, powered by Modulate's homegrown models, which rank #1 on the Hugging Face Open ASR Leaderboard. Powered by Modulate's Ensemble Listening Model (ELM) architecture, the API helps organizations identify costly conversational blind spots - like emotion, emphasis, and frustration signals - that transcript-only monitoring systems frequently miss. Speaking alongside industry luminaries like Geoffrey Hinton and Andrew Ng, Modulate CTO and co-founder Carter Huffman will present a solo talk titled "Voice Agent Supervision: Finding the Hidden Cost of Conversational Blind Spots." During the session, Huffman will explain where traditional AI fails to create expected ROI, and why measuring more than just technical accuracy is necessary to facilitate successful customer interactions. At Ai4, Modulate will demonstrate how its Velma platform provides a voice-native supervision layer for live conversations. By preserving the conversational signals that are often lost when audio is converted to text, Velma enables organizations to identify customer frustration, escalation risk, compliance issues, and other conversational blind spots before they become operational or financial liabilities. This enables enterprises to identify where voice agents may be: * Failing to recognize customer frustration or confusion * Responding inappropriately to urgency, vulnerability, or emotional distress * Repeating information without moving the interaction toward resolution * Missing emphasis that changes the meaning or importance of a customer's request * Making inaccurate claims or commitments * Violating organizational policies or approved communication guidelines * Escalating tension through an unsuitable tone or response * Creating avoidable transfers, callbacks, abandonment, or churn Modulate helps organizations translate these interaction-level failures into operational and financial metrics. Carter Huffman reveals where AI fails to create ROI at Ai4 2026. Modulate CTO and co-founder Carter Huffman will present "Voice Agent Supervision: Finding the Hidden Cost of Conversational Blind Spots" on Wednesday, August 5th from 4:05 to 4:25pm PDT in Palazzo Ballroom D. During the session, Huffman will explain why some of the most expensive voice-agent failures do not appear as traditional software defects. Instead, they emerge when an agent misses tone, emotion, emphasis, hesitation, frustration, or other conversational signals that were never captured in the transcript. "Voice agents can appear successful on a traditional scorecard while quietly creating friction throughout the conversation," said Huffman. "An agent may produce a factually acceptable response and still miss that the customer is confused, losing patience, or emphasizing that the issue is urgent. If enterprises cannot measure those moments, they cannot understand the true performance or ROI of their voice-agent deployments." Attendees will learn how to: * Identify conversational blind spots in voice-agent deployments * Treat tone and emphasis as measurable, first-class signals * Detect and correct friction while a conversation is still underway * Distinguish technical accuracy from genuine conversational understanding * Translate agent-level errors into financial and operational metrics * Build a clearer and more defensible ROI case for enterprise voice AI Huffman is CTO and co-founder of Modulate, where he leads the development of the company's purpose-built Voice Intelligence Engine. A physicist and machine-learning expert trained at MIT, with previous experience at NASA's Jet Propulsion Laboratory, Huffman has spent more than a decade developing AI systems capable of understanding real-world conversational audio beyond the words spoken. Ranked #1 on the Hugging Face Open ASR Leaderboard. At Ai4, Modulate will also highlight its recent recognition as the #1 model on Hugging Face's Open ASR Leaderboard, outperforming models from Microsoft, OpenAI, NVIDIA, IBM, ElevenLabs, AssemblyAI, BosonAI, and other commercial and open-source providers. Modulate ranked first among 88 models evaluated across standardized datasets covering multiple domains, accents, speakers, and recording conditions. The milestone demonstrates that Modulate's voice-native architecture can provide the accurate, low-latency, and cost-efficient transcription foundation required for production voice applications. Modulate's transcription APIs are priced between $0.025 and $0.06 per hour, making them up to 10 times less expensive than several other leading commercial transcription services. Modulate trains its models using more than 500 million hours of noisy, real-world audio. The company developed its technology in demanding, high-scale voice environments where conversations are live, emotionally complex, overlapping, and rarely studio-quality. That foundation now supports enterprise applications across voice-agent supervision, customer experience, fraud prevention, trust and safety, compliance, and conversational intelligence. Meet Modulate at Ai4. Attendees can visit Modulate at Booth 940 to experience live demonstrations of Velma and learn how voice-native conversational intelligence can help enterprises supervise AI agents, improve customer experiences, detect emerging risk, and identify the value being lost through conversational blind spots. Join Modulate for Happy Hour. Modulate and Axonis will co-host a Happy Hour at Ai4 on Tuesday, August 4, from 6:00-8:00 p.m. at SUGARCANE, The Venetian in Las Vegas. Attendees are invited to join both teams for drinks and conversation. Register here: https://luma.com/Axonis-Modulate-Ai4 About Modulate. Modulate is a voice intelligence company building AI models and APIs designed to understand real-world conversational audio at scale. Its technology combines speech recognition, acoustic analysis, and conversational context to deliver reliable, explainable, and cost-effective voice intelligence for developers and enterprises. Media Contact Kristin Canders Grithaus Agency (e) [email protected]