Inworld AI

Inworld AI

Real-time AI NPC engine for gaming

Overview

Inworld AI creates AI-powered non-player characters for games. These NPCs have distinct personalities, contextual awareness, memory, and emotional intelligence to deepen immersion. Their character engine runs in real-time with low latency and includes safety, knowledge, memory, narrative controls, and multimodal abilities, built by the Dialogflow team for production-ready scalability. The product helps game developers monetize by selling access to the engine to enable autonomous, goal-driven NPCs that can learn and adapt within the game world.

About Inworld AI

Simplify's Rating
Why Inworld AI is rated
C+
Rated B on Competitive Edge
Rated C on Growth Potential
Rated C on Differentiation

Industries

Consumer Software

AI & Machine Learning

Entertainment

Gaming

Company Size

51-200

Company Stage

Series A

Total Funding

$117M

Headquarters

Mountain View, California

Founded

2021

Get referred to Inworld AI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • June 2026 Tencent partnership opens enterprise distribution across 3,200-node real-time infrastructure.
  • July 2026 fivefold revenue growth signals demand despite brutal inference economics.
  • Over 18 open roles in July 2026 show active hiring across product and research.

What critics are saying

  • OpenAI, ElevenLabs, and Google pressure pricing; Inworld cut prices over 50%.
  • TTS-2 remains a research preview, risking adoption if reliability lags production needs.
  • Gaming focus stays narrow; one major platform shift can make character engines obsolete.

What makes Inworld AI unique

  • June 2026 Tencent Cloud embedded Inworld TTS into Tencent RTC globally.
  • May 2026 Realtime TTS-2 listens to live audio, not just transcripts.
  • Inworld’s Dialogflow-origin team and top-ranked TTS models give real-time voice depth.

Help us improve and share your feedback! Did you find this helpful?

Funding

Total Funding

$117M

Above

Industry Average

Funded Over

4 Rounds

Series A funding typically happens when a startup has a product and some customers, and now needs funding to scale. This money is usually used to grow the team, expand marketing, and improve the product. Venture capital firms are frequently the main investors here.
Series A Funding Comparison
Above Average

Industry standards

$15M
$8.2M
Discord
$15M
Canva
$30M
Kalshi
$50M
Inworld AI

Benefits

Remote friendly

Flexible work hours

Unlimited PTO

Competitive compensation

Medical, dental, and mental health

Tech setup

Growth & Insights and Company News

Headcount

6 month growth

2%

1 year growth

4%

2 year growth

0%
TRTC
Jun 17th, 2026
Tencent Cloud and Inworld AI announce strategic partnership to deliver a one-stop, lifelike, Realtime Voice AI solution.

Tencent Cloud and Inworld AI announce strategic partnership to deliver a one-stop, lifelike, Realtime Voice AI solution. Jun 17, 2026 Tencent Cloud, the cloud business of global leading technology company Tencent, today announced a strategic partnership with Inworld AI, a leading voice AI research lab and developer platform for real-time, human-level voice experiences. Through the collaboration, Inworld's top-ranked text-to-speech (TTS) models will be deeply integrated with Tencent Cloud's enterprise-grade global realtime communication infrastructure. Developers can directly select Inworld TTS in the Tencent Real-Time Communication (Tencent RTC) console and SDK to build AI applications with emotionally intelligent, real-time voice experiences delivered over Tencent RTC's low-latency global network. Tencent RTC and Inworld AI Power Enterprise-Grade Realtime Voice AI for Global Customers Inworld AI is widely recognized for powering highly expressive voice AI - for companions, enterprise agents, interactive entertainment, and more - where emotional nuance, timing, and non-verbal cues matter as much as the words themselves. Inworld's TTS stack delivers ultra-realistic, context-aware speech and precise voice cloning with performance ranked #1 by users on the Artificial Analysis Speech Arena. Inworld TTS delivers sub-130ms first-chunk latency, supports over 100 languages, and enables instant cross-lingual conversion within a single generation while preserving a consistent speaker voice. Developers can easily clone voices, design custom voices, control vocal style, and stream natural voice responses. By integrating Tencent RTC technology, developers gain access to an enterprise-grade realtime communication backbone featuring more than 3,200 global nodes, sub-300ms worldwide latency. Advanced capabilities, such as AI noise suppression and weak-network resilience, ensure conversational AI applications run seamlessly across regions, especially in some markets where connectivity and latency challenges are common. As part of this partnership, Tencent RTC and Inworld AI have jointly launched a Conversational AI Demo, showcasing the seamless integration of Inworld TTS with Tencent RTC's real-time communication infrastructure. The demo also features a curated selection of recommended voices tailored to different languages and application scenarios, allowing developers to experience them firsthand. Interested developers and enterprises can experience the demo here: https://trtc.io/demo/homepage/#/inworld Wison Xie, Head of Tencent RTC Product at Tencent Cloud, said, "Production-grade voice AI requires both a capable model and a capable network - neither side can deliver it alone. Inworld has set a new bar for expressive, controllable TTS, and we are proud to bring those voices into Tencent RTC's global real-time communication backbone. This partnership is part of a long-term commitment to give developers a one-stop path from prototype to production, so they can focus on the experience they want to create, not the underlying infrastructure." Kylan Gibbs, CEO at Inworld AI, said, "Tencent Cloud gives developers a powerful platform for integrating enterprise-grade conversational AI into any application. Now, with access to Inworld TTS through Tencent RTC, developers can leverage state-of-the-art tech in their full cloud stack. We're excited to partner with Tencent to push the boundaries of high-quality, broadly accessible voice AI." About Tencent Cloud: Tencent Cloud, one of the world's leading cloud companies, is committed to creating innovative solutions to resolve real-world issues and enabling digital transformation for smart industries. Through its extensive global infrastructure, Tencent Cloud provides businesses across the globe with stable and secure industry-leading cloud products and services, leveraging technological advancements such as cloud computing, Big Data analytics, AI, IoT, and network security. It is its constant mission to meet the needs of industries across the board, including the fields of gaming, media and entertainment, finance, healthcare, property, retail, travel, and transportation. About Tencent RTC: Tencent RTC provides real-time communication solutions, including audio/video calling, live streaming, and in-game voice. With enterprise-grade security, AI-powered enhancements, and a global network of over 3,200 nodes, Tencent RTC powers mission-critical communication for customers worldwide. About Inworld AI: Inworld is widely recognized for powering highly expressive voice AI - for companions, enterprise agents, interactive entertainment, and more - where emotional nuance, timing, and non-verbal cues matter as much as the words themselves. As a premiere AI research lab building the models and infrastructure for real time AI experiences, Inworld leads with #1-ranked TTS models optimized for human-like expression and sub-200ms latency that feels like a real conversation. Inworld's Realtime API unifies STT, LLM routing, and TTS into a single low-latency voice pipeline trusted by developers building consumer applications, AI companions, and voice agents worldwide.

Business Insider
Jun 10th, 2026
AI voice startup Inworld cuts pricing by over 50% as inference costs threaten consumer AI startups

Inworld, an AI voice startup that has raised over $117 million, is slashing prices by more than 50% to help consumer AI startups survive escalating infrastructure costs. CEO Kylan Gibbs says inference costs have become the "single number one problem" for consumer AI companies, with many spending 70% to 90% of operating budgets on running AI models. Unlike enterprise software, consumer AI startups typically charge users just $5 to $10 monthly, making profitability difficult as engagement rises. Gibbs notes that whilst users love these products, "every time they become successful, their profitability falls." The company, which reports fivefold revenue growth since early 2026, is betting cheaper AI infrastructure will enable consumer applications in education, therapy and fitness to reach massive scale. Inworld is offering deeper discounts as customers grow.

Inworld AI
May 14th, 2026
Realtime voice agents can now see, listen, and engage.

Realtime voice agents can now see, listen, and engage. Summarize with: Realtime interactions are becoming the primary way people interact with agents. To showcase what's possible with Inworld's latest voice model, Realtime TTS-2, Inworld partnered with Stream to build a flagship reference implementation using their open-source Vision Agents framework. The 'Crashout Buddy' watches your face, hears your words, and shapes its delivery in real time based on how the end user actually feels in the moment. This reference example can be adapted to many more use cases: from professional coaching with realtime guidance, companion apps that notice context and environment, patient intake with verbal and non-verbal cues, 1:1 personalized customer experiences, and beyond. What this means for realtime voice agents. Realtime TTS-2 + Vision Agents only requires a few lines of code to unlock the ability for AI to make users feel heard and stay engaged: Emotional context and conversational awareness become first-class inputs. Whatever your agent knows about the user from sentiment in transcript, signals from a vision model, sensor data, etc. can be turned into a steering directive that shapes how the voice delivers the response. Multilingual deployment prevents a fragmented voice identity. One voice across 100+ languages means a single agent persona for a global user base. No need for model swaps mid-conversation when a user switches languages. Non-verbal sounds become part of the script. Laughs, sighs, breaths, and pauses live inline alongside the words. The agent sounds like it's actually listening before responding. Long conversations build on themselves. Because Realtime TTS-2 carries conversational context forward, sustained interactions stop feeling like a series of disconnected scripts. Sub-200ms latency unlocks interruption and barge-in patterns suitable for realtime agent loops. "We've had early access to Inworld's Realtime TTS-2 for a few days and we're all blown away. The expressiveness, language steering and multi-lingual support are genuinely impressive. The subtle details like natural pausing make it hard to differentiate between AI and human." Neevash Ramdial · Vision Agents Lead, Stream Voice steering, live. The same steering capability powering Crashout Buddy. Pick a delivery tag and hear how the voice changes with the same line. Try a delivery tag [speak tired but warm, like she just got home from a long day]I missed you. How was today? End-of-day affection. Lower energy, gentle smile. Start building. Any use case can follow a simple pattern: take the user's video feed, run a lightweight perception model on it, use the results as context, and let an expressive voice model render the response with appropriate delivery. Vision Agents makes the orchestration simple. Inworld Realtime TTS-2 makes the voice interaction believable. * Vision Agents - the framework * Inworld Realtime TTS-2 - expressive voice * GitHub repo - the code The next generation of agents doesn't read scripts. It understands the full context of the user and makes them feel truly understood.

Business Wire
May 5th, 2026
Inworld launches Realtime TTS-2 voice model giving AI agents contextual empathy in conversations

Inworld AI has launched Realtime TTS-2, a voice model that gives AI agents contextual awareness and emotional intelligence in real-time conversations. The system analyses users' emotional state, tone and pacing to determine not just what to say, but how to say it. Unlike conventional text-to-speech models, TTS-2 processes the user's audio, conversational history and emotional context before generating speech. Developers can control the model using natural language descriptions, with support for over 100 languages whilst preserving voice identity. The model builds on Inworld's TTS-1.5, which ranks first on the Artificial Analysis Speech Arena. Inworld Realtime TTS-2 is now available via the Inworld API, with integration partners including LiveKit, Vapi and Voximplant. The Mountain View-based company has raised over $125 million from investors including Lightspeed Venture Partners and Kleiner Perkins.

The Mirror Democrat and Savanna Times-Journal
May 5th, 2026
Inworld launches new frontier voice model that gives AI agents contextual empathy.

Inworld launches new frontier voice model that gives AI agents contextual empathy. * 10 hrs ago Inworld AI launches Realtime TTS-2, a new generation of voice model built to evolve how AI agents handle realtime conversation. The model is able to understand the full context of conversations and the user's emotional state, tone and pacing to determine not only what to say, but how to say it. The result is voice AI that feels as good as it sounds. Realtime conversation is the most human way to connect with each other. Now, as AI takes up more of its conversations, it must evolve beyond intelligence to afford the same shared emotionality and context-awareness that makes connection meaningful. But until today, voice AI has been tuned for static audiobooks and voiceovers rather than live connection. It has been text-to-speech and speech-to-text in the most literal sense: words converted to audio, and audio converted to words, making AI voice interactions feel more like a misinterpreted text message than a meaningful conversation. The problem with today's voice AI When a frustrated customer calls support, today's voice agents respond with the same bright, even tone they use for everything, because they have no awareness of how the caller is speaking. Inworld's Realtime TTS-2 hears the frustration. Its voice softens and its pace slows. It reasons through the weight of the moment before it responds. Or consider a patient calling to discuss lab results. They start the call measured and calm. Then the agent shares an unexpected finding. The patient's voice tightens; their questions come faster. Today's voice agents would barrel ahead at the same pace and pitch, ignoring the gravity of the situation. TTS-2 registers the shift in real time. It slows down. It leaves space. It delivers the next piece of information with steadiness and care, not because someone scripted a "nervous patient" pathway, but because the model heard how the person was speaking and adapted the way a human would. The reason today's voice agents sound mechanical in conversation is architectural. Conventional voice models receive a string of text and produce audio. They have no access to how the user sounds, what was said before, or what the moment requires. A new kind of voice model Inworld's previous model, TTS-1.5, already ranks #1 on the Artificial Analysis Speech Arena, above Google and ElevenLabs. With voice quality achieved, Inworld set out to build TTS-2 with a fundamentally different architecture - one that can process conversation the way a human listener would, before a single word is spoken. Before speech is generated, TTS-2 captures the user's audio and extracts context, emotion, and tone in real time. It then reasons over the full conversational history: what was said in previous turns, what the most important moments were, what can be inferred from how the user sounds right now. From this, it estimates the user's emotional state and determines the agent's appropriate response state: not just what to say, but how to say it. What non-verbal expressions are appropriate. How what the agent says might land given everything that came before. Inside TTS-2, all of that context converges. The model receives what to say, how to say it (full natural-language voice direction, not preset emotion tags), direct audio from the user (to further condition expression in real time), and the full conversation history. TTS-2 synthesizes all of this into emotionally aware, contextual speech, adjusting tone, pacing, and delivery based on the complete picture of the interaction. The result is a voice system that sounds like a person in conversation, not a person reading an audiobook. Developers steer the model with natural language the way they prompt an LLM: full descriptions like [act like you just got home from a long day, tired but warm], combined with inline controls for specific moments ([whispering], [sigh], [excited]). The voice is as controllable as it is expressive, across over 100 languages with on-the-fly switching inside a single generation, preserving the speaker's voice identity across every language. "We are obsessed with how voice AI feels, not just how it sounds. Realtime voice is the most natural way for people to communicate with AI, because it is the most natural way people communicate with each other. Voice is how we actually connect. We built TTS-2 to make that connection feel real," said Kylan Gibbs, CEO and Co-Founder of Inworld AI. "Most TTS models generate speech in isolation from the conversation around them. TTS-2 is trained to use audio context from the full multi-turn exchange, and take voice direction so how the model speaks adjusts to how it was spoken to. Building a system that does this in real-time, at production quality, with full controllability, required solving problems that the field had treated as future work for years. It is a different generation of system than a text-to-audio model, and it is what is required for voice AI that behaves naturally inside a realtime pipeline," said Igor Poletaev, Chief Science Officer at Inworld AI. Availability Inworld Realtime TTS-2 is available via the Inworld API, and as part of the Inworld Realtime API for end-to-end speech-to-speech over a single persistent connection. Integration partners include Layercode, LiveKit, NLX, Pipecat, Vapi, and Voximplant. Developers can try the live demo or learn more at inworld.ai/tts. See inworld.ai/pricing for current rates. About Inworld AI Inworld is a research lab focused on solving realtime interaction. The company's Realtime TTS is ranked #1 on the Artificial Analysis Speech Arena, with two of the top five positions. Realtime STT offers speech recognition that includes voice profiling to detect detailed user context. The Realtime Router is a user-aware reasoning layer that selects the optimal model and prompt for every context. And Realtime API unifies everything into a single persistent connection for full-duplex conversational AI that includes contextual understanding with natural speech. The founding team comes from DeepMind and Google, and they have raised $125M+ from leading investors like Lightspeed Venture Partners, Section 32, Bitkraft, Kleiner Perkins, and Founders Fund. Media gallery

Recently Posted Jobs

Sign up to get curated job recommendations

Inworld AI is Hiring for 14 Jobs on Simplify!

Find jobs on Simplify and start your career today

Don't see your dream role? Check out thousands of other roles on Simplify. Browse all jobs →