Hume AI

Hume AI

Offers emotion-aware AI datasets and models

Overview

Hume AI provides large-scale training datasets and AI models to help build systems that understand and respond to human emotions. It focuses on integrating empathy into AI across applications such as social networks, digital assistants, health tech, and education. The product works by offering scientifically-backed datasets and models that can detect and respond to users' emotional states, grounded in research on over 30 distinct emotions. This makes it possible for developers to create emotionally aware products that improve user experience and well-being. The company differentiates itself with a strong emphasis on empirical emotional science and empathy-driven capabilities, rather than generic AI tools. Its goal is to ensure AI serves human goals and emotional well-being by providing tools that enable empathic technology.

About Hume AI

Simplify's Rating
Why Hume AI is rated
B
Rated B on Competitive Edge
Rated A on Growth Potential
Rated C on Differentiation

Industries

Data & Analytics

Enterprise Software

AI & Machine Learning

Healthcare

Company Size

51-200

Company Stage

Series B

Total Funding

$75.7M

Headquarters

New York City, New York

Founded

2021

Get referred to Hume AI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Hume launched TADA in March 2026, delivering zero hallucinations across 1,000-plus samples.
  • Real World VoiceEQ in July 2026 highlights benchmark leadership in emotion, listening, and robustness.
  • Andrew Ettinger said January 2026 revenue is tracking above $100 million from research partnerships.

What critics are saying

  • Google hired Alan Cowen and seven engineers in January 2026, weakening execution depth.
  • ElevenLabs, OpenAI, and Google Gemini now compete directly on voices, cloning, and language coverage.
  • Open-source TADA lowers switching costs; if labs internalize stack pieces, Hume's revenue shrinks.

What makes Hume AI unique

  • Emotion-sensing voice tech remains Hume's moat, not generic TTS.
  • Hume's research infrastructure spans data, evaluation, and reinforcement-learning workflows for frontier labs.
  • The January 2026 Google licensing deal validates Hume's emotional voice datasets and models.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$75.7M

Above

Industry Average

Funded Over

5 Rounds

Series B funding is typically for startups that have proven their business model and need more funding to expand rapidly—often by entering new markets or adding more products. Investors are usually venture capital firms that specialize in later-stage investments.
Series B Funding Comparison
Above Average

Industry standards

$35M
$45M
Linktree
$50M
Hume AI
$65M
Substack
$100M
ClickUp

Benefits

Remote Work Options

Growth & Insights and Company News

Headcount

6 month growth

-6%

1 year growth

1%

2 year growth

-3%
Andres SEO Expert LLC
Jul 15th, 2026
HumeAI's Real World VoiceEQ benchmark exposes Voice AI's blind spots in human interaction.

HumeAI's Real World VoiceEQ benchmark exposes Voice AI's blind spots in human interaction. What are You Looking For? July 15, 2026 HumeAI's Real World VoiceEQ benchmark reveals voice AI excels at speaking but fails at listening, highlighting evaluation gaps. Visualizing the gap in sound wave transmission highlighted by HumeAI's VoiceEQ benchmark. By Andres SEO Expert. Key takeaways. * Voice AI progress is specialized; no single model excels in all capabilities. * Models are stronger at speaking than listening, missing critical paralinguistic cues. * Traditional benchmarks overestimate real-world performance, especially in noisy environments. * Human evaluation remains irreplaceable for assessing subjective voice quality. Beyond WER: Real World VoiceEQ benchmarks the soul of Voice AI. Today, HumeAI unveiled Real World VoiceEQ, a comprehensive benchmark designed to measure the human quality of voice interactions. Unlike traditional metrics focused on word error rate and latency, this benchmark evaluates how well voice AI systems recognize emotion, tone, and context - the invisible layers of human communication that transcripts miss. Built from over 1 million human ratings across diverse demographics and environments, it ranks more than 40 leading voice models across 15+ evaluation dimensions. Table of contents. The benchmark: methodology and metrics. Real World VoiceEQ evaluates both proprietary and open-source voice models across automatic speech recognition (ASR), text-to-speech (TTS), speech-to-speech (S2S), and speech understanding tasks. The benchmark includes 785,000 TTS ratings and 48,000 STS ratings, making it one of the largest human evaluations of voice AI conducted to date. All evaluations were conducted using Kairos, HumeAI's voice-native evaluation platform. The same infrastructure allows frontier AI labs and enterprises to run custom evaluations, identify failure modes, generate human preference data, and improve models via reinforcement learning. HumeAI's Real World VoiceEQ benchmark covers more than 60 metrics across 15+ dimensions, including emotional understanding, speaker consistency, expressiveness, robustness to noise and accents, and conversational intelligence. By breaking down performance into specialized capabilities, HumeAI aims to provide a more nuanced view of voice model quality than traditional aggregate scores. Key findings: specialization, listening gap, and benchmark limitations. Specialized progress: no single best model. HumeAI's results reveal that voice AI progress is becoming increasingly specialized. No system configuration ranked among the top five across all eight capability groups in TTS evaluations. One model may excel at precise pronunciation for technical terms but fail to produce emotionally expressive speech, while another sounds natural but struggles with accuracy. This underscores the importance of evaluating capabilities independently rather than collapsing them into a single score. Speaking vs. Listening: A critical gap. Speech-to-Speech models showed the widest variation. While some systems recognized emotion well, they often failed to respond appropriately. HumeAI found that many models remain largely transcript-driven, ignoring paralinguistic cues such as tone, pacing, hesitation, and emphasis. These cues are critical for interpreting confidence, uncertainty, sarcasm, and empathy in real conversation. For instance, a confident 'Yes' versus a hesitant '...yes...' have identical transcripts but vastly different meanings - yet most current voice AI cannot distinguish them. Benchmark limitations: overestimation and real-world failure. Traditional benchmarks near saturation do not reflect real-world conditions. HumeAI observed that word error rates on noise-backed speech were roughly four times higher than on music-backed speech, showing how a single background-audio score can hide true failure modes. Moreover, the research indicates that some models may be over-optimized for established benchmarks, even reproducing known errors in reference transcripts. Human evaluation remains essential. When comparing leading speech-language models (SLMs) with trained human raters, agreement was highest on objective tasks like pronunciation accuracy but dropped significantly on subjective judgments such as whether a voice fit a role or maintained consistent identity. Automated evaluators are not yet a substitute for human listeners for tasks requiring acoustic perception and social interpretation. Strategic implications: human-grounded metrics and industry trends. The launch of Real World VoiceEQ comes at a critical juncture for the voice AI industry. As voice interfaces become dominant in customer support, healthcare, education, and personal assistants, the inability of current models to fully understand human communication poses a bottleneck. HumeAI's benchmark provides a necessary corrective to the industry's over-reliance on narrow technical metrics. This focus on human quality aligns with broader trends in AI efficiency and specialization. Recent work from Cohere on hardware-aware dynamic speculation demonstrated a 23% inference speedup for large language models, emphasizing the need for efficient deployment without sacrificing quality. Similarly, NVIDIA's one-day vision model post-training approach achieved 93% accuracy by leveraging agent skills, illustrating how specialized techniques can rapidly improve model capabilities. These developments, coupled with HumeAI's findings, suggest that the future of voice AI lies in a combination of efficient architectures, specialized models, and human-centered evaluation frameworks. The recognition that voice AI systems must listen as well as they speak will drive demand for richer training data, multi-modal understanding, and reinforcement learning from human feedback. Companies that invest in these areas will have a competitive advantage as users increasingly expect natural, empathetic interactions. Conclusion: the path to truly human Voice AI. Real World VoiceEQ marks a significant step toward measuring what truly matters in voice AI: the ability to communicate like a human. By exposing the gaps between benchmark performance and real-world experience, HumeAI challenges the industry to broaden its definition of success. Speed and accuracy are table stakes; understanding and expression are the differentiators. As voice becomes AI's primary interface, the models that succeed will be those that can listen, adapt, and respond with genuine human quality. HumeAI's benchmark provides the tools to measure that progress, and the insights to guide it. Staying ahead in the rapidly shifting landscape of AI requires precision. To future-proof your digital strategy and scale effortlessly, you need a foundation built on precision. Optimize your site with advanced speed engineering, secure your infrastructure in high-performance hosting environments, and streamline your entire workflow through autonomous AI pipelines. If you are ready to elevate your systems, Connect with Andres at Andres SEO Expert to build your ultimate architecture. Frequently asked questions. What is Real World VoiceEQ? How is Real World VoiceEQ different from traditional voice AI benchmarks? What are the key findings of the Real World VoiceEQ benchmark? Why is there a gap between speaking and listening in voice AI? What are the limitations of current voice AI benchmarks highlighted by HumeAI? What does HumeAI's benchmark mean for the future of voice AI? How can voice AI improve human-like communication? July 15, 2026

SearchYour.AI
Mar 10th, 2026
Hume AI launches TADA, a fast open-source voice system that eliminates hallucinations

Hume AI launches TADA, a fast open-source voice system that eliminates hallucinations. 10/03/2026 Hume AI releases TADA under an open-source license, a text-to-speech system that synchronizes text and audio to eliminate content errors and achieve five times the speed of current systems. Hume AI has released TADA (Text-Acoustic Dual Alignment), a voice generation system that addresses one of the most common problems in current large language model-based systems: the mismatch between how text and audio are represented. Conventional text-to-speech systems generate between 12.5 and 75 acoustic signal frames per second of audio, compared to just 2 or 3 text tokens. This gap forces models to handle very long sequences, which slows down processing and increases the risk of the system skipping words or inserting non-existent content - a flaw known as hallucination. TADA resolves this imbalance with a tokenization scheme that assigns exactly one continuous acoustic vector per text token. As a result, text and audio are processed in parallel and at the same rate, without compressing the audio or adding extra intermediate layers. In terms of speed, the system achieves a real-time factor of 0.09 - more than five times faster than comparable LLM-based text-to-speech systems. In tests with over 1,000 samples from the LibriTTSR dataset, the model produced zero hallucinations. In human evaluations on expressive, long-form speech, it scored 4.18 out of 5 for speaker similarity and 3.78 out of 5 for naturalness, ranking second overall. The model's compact size allows it to run on mobile devices without relying on cloud services. In terms of context management, it can handle up to 700 seconds of audio within a 2,048-token context window, compared to around 70 seconds for conventional systems under the same conditions. Hume AI is releasing two versions: a one-billion-parameter model for English and a three-billion-parameter multilingual model supporting eight languages. Both are available on Hugging Face under an open-source license. The researchers themselves acknowledge limitations still to be resolved, including potential speaker drift during very long generations and reduced text quality when generating text and speech simultaneously. Key points. * TADA is a new open-source text-to-speech system developed by Hume AI. * It synchronizes text and audio in a 1:1 ratio, eliminating the mismatch in current systems. * It is more than five times faster than comparable LLM-based TTS systems. * In tests with over 1,000 samples, it produced zero hallucinations. * It is lightweight enough to run on mobile devices without a cloud connection. * It can handle up to 700 seconds of audio versus 70 seconds in conventional systems. * Available in two versions: 1B parameters in English and 3B multilingual across eight languages. * Still has limitations in very long generations and when combining text and speech simultaneously. Videos. Links. Related AI. Interfaz de voz con inteligencia emocional. Research laboratory and technology company specialized in AI models with emotional intelligence. Its main model integrates voice and language processing, with adjustable voice synthesis in timbre,... Lastest news. * Copilot Cowork, the AI agent that manages and executes tasks within Microsoft 365 09/03/2026 Microsoft has announced Copilot Cowork, a new Microsoft 365 Copilot feature that goes beyond chat to execute complete tasks autonomously,... * OpenAI launches GPT-5.4, its most powerful AI model for professional work 05/03/2026 OpenAI has released GPT-5.4, a model that combines advanced reasoning, coding and computer control into a single tool designed for complex... * Anthropic refuses to back down against the Pentagon while OpenAI signs an alternative deal 02/03/2026 Anthropic rejects removing two restrictions on its AI's use by the military, in a conflict that led Trump to order its removal from all federal... * Perplexity Computer, the system that coordinates multiple AI models at once 25/02/2026 Perplexity introduces Computer, an AI agent capable of creating and executing complete workflows for hours or months, autonomously coordinating...

PR Newswire
Jan 22nd, 2026
Hume AI names new CEO, targets $100M revenue from voice AI research partnerships

Hume AI, a voice AI research company, has appointed Andrew Ettinger as chief executive officer. Ettinger brings 15 years of experience in data and AI infrastructure, having built teams responsible for over $2 billion in annual recurring revenue at Pivotal, Astronomer and Appen, where he most recently served as chief revenue officer. The company has also agreed to non-exclusively licence certain technologies to Google, with co-founder Alan Cowen joining Google. Hume has been expanding its research platform, including voice evaluation tools, data pipelines and reinforcement-learning infrastructure for training voice AI models. Ettinger said Hume is on track to generate more than $100 million in revenue this year from research partnerships with frontier AI labs and enterprises. The company plans to release the next generation of its text-to-speech and speech-to-speech models shortly.

Wired
Jan 22nd, 2026
Google acquires Hume AI CEO and engineers in licensing deal to boost voice emotion tech

Google DeepMind has hired Hume AI's CEO Alan Cowen and approximately seven engineers as part of a licensing agreement with the AI voice startup. Financial terms were not disclosed, though Hume AI will continue supplying its technology to other AI labs. Hume AI specialises in emotionally intelligent voice interfaces, training models to detect and respond to emotional cues in users' voices. The company has raised $74 million and expects $100 million in revenue by 2026. Cowen, who holds a PhD in psychology, will help Google integrate voice and emotional intelligence into its frontier models. The deal positions Google to compete more aggressively with OpenAI's ChatGPT voice mode and follows Google's partnership with Apple to power Siri with Gemini. The arrangement represents another talent acquisition that avoids traditional merger oversight.

Parler
Oct 16th, 2025
Niantic's Peridot, the Augmented Reality Alien Dog, Is Now a Talking Tour Guide | WIRED

Niantic's Peridot, the augmented reality alien dog, is now a talking tour guide | WIRED. Niantic is giving its cute AR cartoon companions a voice that will let them guide you around in the real world and point out interesting facts. The feature is being demo'd first in Snap Spectacles. Imagine you're walking your dog. It interacts with the world around you - sniffing some things, relieving itself on others. You walk down the Embarcadero in San Francisco on a bright sunny day, and you see the Ferry Building in the distance as you look out into the bay. Your dog turns to you, looks you in the eye, and says, "Did you know this waterfront was blocked by piers and a freeway for 100 years?" OK now imagine your dog looks like an alien and only you can see it. That's the vision for a new capability created for the Niantic Labs AR experience Peridot. Niantic, also the developer of the worldwide AR behemoth Pokémon Go, hopes to build out its vision of extending the metaverse into the real world by giving people the means to augment the space around them with digital artifacts. Peridot is a mobile game that lets users customize and interact with their own little Dots - dog-sized digital companions that appear on your phone's screen and can look like they're interacting with the world objects in the view of your camera lens. They're very cute, and yes, they look a lot like Pokémon. Now, they can talk. Peridot started as a mobile game in 2022, then got infused with generative AI features. The game has since moved into the hands of Niantic Spatial, a startup created in April that aims to turn geospatial data into an accessible playground for its AR ambitions. Now called Peridot Beyond, it has been enabled in Snap's Spectacles. Hume AI, a startup running a large language model that aims to make chatbots seem more empathetic, is now partnering with Niantic Spatial to bring a voice to the Dots on Snap's Spectacles. The move was initially announced in September, but now it's ready for the public and will be demonstrated at Snap's Lens Fest developer event this week.

Recently Posted Jobs

Sign up to get curated job recommendations

Hume AI is Hiring for 4 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs →