Full-Time

Technical Account Manager

Cartesia

Cartesia

51-200 employees

Develops foundation models with subquadratic architectures

Compensation Overview

$220k - $280k/yr

+ Equity package + Commuter allowance

H1B Sponsorship Available

San Francisco, CA, USA

In Person

In-office work is required in San Francisco.

Category
Sales & Account Management (1)
Required Skills
LLM
REST APIs

Get referred to Cartesia

See people who can refer or advise you

Requirements
  • At least 6 years of experience in technical account management, customer success, solutions engineering, or a related customer-facing technical role at a high-growth business-to-business company.
  • Ability to read application programming interface documentation, understand system architecture, triage integration issues, and hold substantive technical conversations with engineers.
  • Ability to manage a complex, multi-stakeholder account book simultaneously and prioritize effectively.
  • Ability to run complex implementations end-to-end while keeping multiple accounts moving.
  • Ability to scope customer problems clearly, route them correctly, and drive them to resolution.
  • Ability to communicate effectively with engineers and vice presidents while preserving substantive detail.
  • Commercial awareness of the connection between customer health and revenue growth, including consumption, retention, and expansion.
  • High ownership and autonomy, with accountability for customer outcomes.
  • Fluency with artificial intelligence agents and tools for managing account workload and automating status tracking.
Responsibilities
  • Own the full post-sale technical relationship for a book of enterprise accounts from pre-go-live scoping through post-launch stabilization and expansion.
  • Drive customers to go-live by defining success criteria with senior technical stakeholders, designing customer-specific implementation strategies, running daily standups and 30/60/90 onboarding reviews, and project-managing rollouts end-to-end.
  • Independently triage customer technical issues by distinguishing product bugs, integration failures, and customer-side misconfigurations, then routing issues to the appropriate internal owner.
  • Lead field development engineers through implementation by setting direction, managing scope, and serving as the primary technical counterpart for the customer.
  • Translate customer needs into clear internal documentation, including bug reports, contextualized feature requests, and scope-change documents.
  • Monitor account health proactively, identify at-risk customers, and drive consumption and usage growth across the account book.
  • Surface expansion opportunities to account managers with specific context, flag upsell-ready accounts, build handoffs, and connect customer health with growth.
  • Build the technical account management playbook through execution, including a 30/60/90 framework, escalation paths, go-live checklists, and technical champion management.
  • Partner with Product and Engineering as the voice of the customer, turn customer pain into platform improvements, and close the feedback loop.
Desired Qualifications
  • Experience in voice artificial intelligence, real-time audio/video, or telephony infrastructure.
  • A background where technical account management was a revenue-driving function rather than a support function.
  • Infrastructure, developer tools, or application programming interface-first product experience.
  • Experience scaling from Series A or B through growth stage and adapting the customer motion as the business evolves.

Cartesia.ai develops advanced AI foundation models using new subquadratic architectures and state space models. Its models are designed as underlying systems that can be adapted for many applications and are released in part under open-source licenses (e.g., Apache 2.0), with licensing, partnerships, and consulting as revenue streams. The company’s products work by providing large, adaptable AI models that clients can license or customize for specific needs, enabling faster processing and efficient inference compared to traditional designs. Cartesia differentiates itself from competitors by using subquadratic, state space approaches instead of relying mainly on standard Transformer models, and by combining open-source releases with enterprise partnerships. The goal is to help businesses and research institutions access powerful, adaptable AI capabilities while building a community around its technology and sustaining revenue through licensing and services.

Company Size

51-200

Company Stage

Late Stage VC

Total Funding

$191M

Headquarters

San Francisco, California

Founded

2023

Get referred to Cartesia

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • June 2026 Sonic-3.5 and Ink-2 launches expanded Cartesia's full-stack real-time voice moat.
  • Cartesia's 42-language stack targets global support, tutoring, and transcription workflows.
  • The $100 million backing from Kleiner Perkins, Index, Lightspeed, and NVIDIA funds scaling.

What critics are saying

  • OpenAI and ElevenLabs pressure pricing, quality, and developer mindshare throughout 2026.
  • Closed roles in June 2026 signal hiring caution after rapid product expansion.
  • Voice AI customers churn fast if latency slips; one bad release kills enterprise trust.

What makes Cartesia unique

  • Cartesia tops streaming TTS and STT leaderboards with Sonic-3.5 and Ink-2.
  • SSM architecture delivers sub-300ms voice loops without Transformer-style full-context replay.
  • ServiceNow integrates Cartesia for secure enterprise voice agents, validating production readiness.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

401(k) Company Match

Relocation Assistance

Growth & Insights and Company News

Headcount

6 month growth

-3%

1 year growth

-8%

2 year growth

0%
PR Newswire
Jul 31st, 2026
Palabra.ai takes #1 spot for speed in new text-to-speech (TTS) benchmark.

Palabra.ai takes #1 spot for speed in new text-to-speech (TTS) benchmark. PALABRA.AI LTD Jul 31, 2026, 10:41 ET Palabra ranks #1 for latency on Coval's independent, open-source text-to-speech benchmark, posting 104 milliseconds - roughly twice as fast as the nearest competitor - with a 6% word error rate. Coval's independent benchmark puts Palabra's text-to-speech model at 104 ms latency - about twice as fast as the next-closest competitors. LONDON, July 31, 2026 /PRNewswire-PRWeb/ - Palabra, a real-time speech AI company, has ranked #1 for latency on Coval's independent TTS benchmark, a widely watched, continuously updated leaderboard tracking how leading voice AI models perform under real-world conditions. Palabra posted 104 milliseconds of latency and a 6% word error rate (WER), outpacing ElevenLabs, Cartesia, and other major players in the space - roughly twice as fast as the nearest competitor. Latency is one of the key factors determining whether a voice AI system can support a fluid, real-time conversation. Even when the generated voice itself sounds highly natural, a noticeable delay before each response can disrupt the flow of dialogue, creating awkward pauses, causing people to talk over each other, and making the interaction feel less immediate. Many voice AI systems still operate at around 500 milliseconds of latency or more. Palabra's model achieves 104 milliseconds in Coval's independent benchmark, significantly reducing the delay between turns and enabling a more fluid, uninterrupted dialogue. Coval's benchmark is built to reflect production reality rather than idealized lab conditions: it measures Time to First Audio (TTFA) - the delay a listener actually notices, including any silence before the first audible sample - using pinned, versioned datasets so every provider is scored against identical inputs. Word error rate is calculated by transcribing each provider's synthesized audio with a fixed ASR model and scoring it against the original text, so the number reflects whether the speech is intelligible, not just fast. The full runner and methodology are open-source and independently reproducible. "Latency is one of the most consequential metrics in voice AI," said Brooke Hopkins, Founder and CEO of Coval. "Humans respond in around 400 milliseconds, and systems slower than that feel unnatural regardless of output quality. What makes Palabra's result significant is what it creates downstream: when a model is this fast, every other component in the stack gets more time to think." Palabra says the result reflects a new TTS architecture that achieves 35ms of time-to-first-audio before network overhead, paired with production infrastructure built to hold that speed at scale across large volumes of concurrent users. "The biggest unsolved challenge that the world's leading speech AI labs are working on today is making interactions with AI voice agents feel indistinguishable from conversations with another person," said Artem Kukharenko, CEO & Co-founder of Palabra Al. "Latency above 200 milliseconds introduces a noticeable delay that disrupts the natural flow of communication. That's why we've focused on building TTS with extremely low time-to-first-audio that could work at large scale in production systems. In Coval's independent benchmark, we rank #1 at 104ms - the closest competitor is approximately twice as slow." The TTS result is part of a broader focus at Palabra on real-time speech models - spanning text-to-speech, speech recognition (ASR), and speech-to-speech translation - built around low latency as a core design principle across the stack. About Palabra Palabra is a voice AI lab developing real-time models for text-to-speech (TTS), automatic speech recognition (ASR), and speech-to-speech translation. Its proprietary models and APIs enable developers and enterprises to build low-latency multilingual voice experiences across customer support, conferencing, live streaming, education, and other real-time applications. Palabra recently raised $8.4 million in pre-seed funding led by Seven Seven Six (776) with participation from Creator Ventures, and prominent angel investors including Max Mullen, co-founder of Instacart; Anne Lee Skates, former partner at Andreessen Horowitz; Mehdi Ghissassi,, former Head of Product at DeepMind; and Namat Bahram, an early backer of ElevenLabs. About Coval Coval is the San Francisco-based evaluation platform for voice AI, founded by ex-Waymo engineer Brooke Hopkins. It builds independent, continuously updated, open-source benchmarks that measure how voice AI systems perform in production, rather than in idealized lab conditions. Coval recently raised a $28M Series A led by Norwest, with Twilio Ventures and Y Combinator also on the cap table, and is already trusted by over 60 organizations, including Zoom and Deepgram. Its public TTS leaderboard is available at benchmarks.coval.ai/tts. Media Contact Dmytro Tymoshenko, PALABRA.AI LTD, 380 939944325, [email protected], palabra.ai SOURCE PALABRA.AI LTD

Vialynx Inc
Jul 31st, 2026
Palabra.ai takes #1 spot for speed in new text-to-speech (TTS) benchmark.

Palabra.ai takes #1 spot for speed in new text-to-speech (TTS) benchmark. Business · JUL 31, 2026 PR Newswire Palabra ranks #1 for latency on Coval's independent, open-source text-to-speech benchmark, posting 104 milliseconds - roughly twice as fast as the nearest... Palabra ranks #1 for latency on Coval's independent, open-source text-to-speech benchmark, posting 104 milliseconds - roughly twice as fast as the nearest competitor - with a 6% word error rate. Coval's independent benchmark puts Palabra's text-to-speech model at 104 ms latency - about twice as fast as the next-closest competitors. LONDON, July 31, 2026 /PRNewswire-PRWeb/ - Palabra, a real-time speech AI company, has ranked #1 for latency on Coval's independent TTS benchmark, a widely watched, continuously updated leaderboard tracking how leading voice AI models perform under real-world conditions. Palabra posted 104 milliseconds of latency and a 6% word error rate (WER), outpacing ElevenLabs, Cartesia, and other major players in the space - roughly twice as fast as the nearest competitor. The hardest problem in speech AI is making voice agents feel indistinguishable from a real person. Latency above 200 milliseconds breaks that. We built our TTS for extremely low time-to-first-audio at production scale. Coval's benchmark ranks us #1 at 104 ms, nearly twice as fast as the next model. Latency is one of the key factors determining whether a voice AI system can support a fluid, real-time conversation. Even when the generated voice itself sounds highly natural, a noticeable delay before each response can disrupt the flow of dialogue, creating awkward pauses, causing people to talk over each other, and making the interaction feel less immediate. Many voice AI systems still operate at around 500 milliseconds of latency or more. Palabra's model achieves 104 milliseconds in Coval's independent benchmark, significantly reducing the delay between turns and enabling a more fluid, uninterrupted dialogue. Coval's benchmark is built to reflect production reality rather than idealized lab conditions: it measures Time to First Audio (TTFA) - the delay a listener actually notices, including any silence before the first audible sample - using pinned, versioned datasets so every provider is scored against identical inputs. Word error rate is calculated by transcribing each provider's synthesized audio with a fixed ASR model and scoring it against the original text, so the number reflects whether the speech is intelligible, not just fast. The full runner and methodology are open-source and independently reproducible. "Latency is one of the most consequential metrics in voice AI," said Brooke Hopkins, Founder and CEO of Coval. "Humans respond in around 400 milliseconds, and systems slower than that feel unnatural regardless of output quality. What makes Palabra's result significant is what it creates downstream: when a model is this fast, every other component in the stack gets more time to think." Palabra says the result reflects a new TTS architecture that achieves 35ms of time-to-first-audio before network overhead, paired with production infrastructure built to hold that speed at scale across large volumes of concurrent users. "The biggest unsolved challenge that the world's leading speech AI labs are working on today is making interactions with AI voice agents feel indistinguishable from conversations with another person," said Artem Kukharenko, CEO & Co-founder of Palabra Al. "Latency above 200 milliseconds introduces a noticeable delay that disrupts the natural flow of communication. That's why we've focused on building TTS with extremely low time-to-first-audio that could work at large scale in production systems. In Coval's independent benchmark, we rank #1 at 104ms - the closest competitor is approximately twice as slow." The TTS result is part of a broader focus at Palabra on real-time speech models - spanning text-to-speech, speech recognition (ASR), and speech-to-speech translation - built around low latency as a core design principle across the stack. About Palabra Palabra is a voice AI lab developing real-time models for text-to-speech (TTS), automatic speech recognition (ASR), and speech-to-speech translation. Its proprietary models and APIs enable developers and enterprises to build low-latency multilingual voice experiences across customer support, conferencing, live streaming, education, and other real-time applications. Palabra recently raised $8.4 million in pre-seed funding led by Seven Seven Six (776) with participation from Creator Ventures, and prominent angel investors including Max Mullen, co-founder of Instacart; Anne Lee Skates, former partner at Andreessen Horowitz; Mehdi Ghissassi,, former Head of Product at DeepMind; and Namat Bahram, an early backer of ElevenLabs. About Coval Coval is the San Francisco-based evaluation platform for voice AI, founded by ex-Waymo engineer Brooke Hopkins. It builds independent, continuously updated, open-source benchmarks that measure how voice AI systems perform in production, rather than in idealized lab conditions. Coval recently raised a $28M Series A led by Norwest, with Twilio Ventures and Y Combinator also on the cap table, and is already trusted by over 60 organizations, including Zoom and Deepgram. Its public TTS leaderboard is available at benchmarks.coval.ai/tts. Media Contact Dmytro Tymoshenko, PALABRA.AI LTD, 380 939944325, [email protected], palabra.ai SOURCE PALABRA.AI LTD Published by News Desk · WeeklyReviewer The WeeklyReviewer news desk monitors breaking developments around the clock - sourcing, verifying, and publishing real-time industry news across business, technology, politics, science, sports, and world affairs. Every story that comes through the Live Wire is reviewed for accuracy before publication, keeping our readers ahead of the curve without the noise.

Baseten
Jul 29th, 2026
Announcing Baseten for Model Labs.

Announcing Baseten for Model Labs. Baseten is excited to launch a new platform designed specifically for model labs. Last updated. July 29, 2026 Baseten is excited to launch Baseten for Model Labs: a set of products and services designed to help closed-weight model labs distribute and monetize their models quickly and easily. Baseten for Model Labs is the central hub for model consumers to access models and for model builders to distribute them. On May 6, Baseten introduced the Frontier Gateway, a managed inference gateway built to give closed-weight model labs a way to serve their model in production using a white-labeled API. Since then, many labs like Poolside, Subconscious, Trajectory, and WRITER have successfully leveraged the gateway to monetize their models. As these labs gained traction with developers, Baseten quickly saw a different need emerge. Developers wanted access to these specialized and powerful models, but they didn't want to sign up for a new API or onboard another sub-processor to use a new model in production. They wanted to access it from their preferred inference platform easily. At the same time, its lab partners saw the opportunity to list their models on the Baseten Model Library as a compelling new distribution channel beyond the white-labeled API Baseten provide them with the Frontier Gateway. Seeing the strong demand from both sides is what led Baseten to collaborate with many labs to build a platform that addresses the needs of both model labs and their downstream customers, which Baseten is excited to launch today. A new distribution platform for closed models. With Baseten for Model Labs, Baseten is expanding the services Baseten offer labs with a new set of capabilities built to help them easily monetize and scale. Baseten built a distribution platform for labs that want to bring their model to production but don't want to stand up the technical and operational infrastructure to do so: * Inference infrastructure built for production out of the box: Model labs don't need to spend months building billing, remittance, API keys, compliance authentication, authorization, compute procurement, regional support and more, to bring their model to production. Baseten handles the infrastructure and operational complexity. * Increased visibility with a highly qualified developer audience: Baseten is well-known by AI developers around the world as the fastest, most reliable, and scalable inference provider. Labs benefit from increased visibility with Baseten's growing developer community and enterprise customers, expanding awareness among highly qualified AI practitioners. * Strong IP protection for labs' closed-weights: Lab model weights available through the Model Library are protected by its strict distribution agreement and secure infrastructure, so end-customers can only consume the model and can't modify it or extract its weights. * GTM Support: When a model is published on the Model Library, model labs join its growing ecosystem and benefit from an awareness boost through joint marketing and co-selling engagements with Baseten's sales team when their model is a strong fit for a customer's need. In short, model labs can get a direct line to monetize their model without having to build the infrastructure, compliance, and GTM motion themselves. Labs that are already using Baseten for Labs. Baseten is proud to launch alongside 15 lab partners and many users who are already leveraging these models for a wide range of use cases: * Cartesia builds state-space models (SSMs) for real-time voice intelligence, featuring Sonic (TTS) and Ink (STT). Replacing traditional transformers for hyper-efficiency, it delivers sub-90ms latency across cloud, edge, or on-premise deployments. Its top use cases include real-time conversational AI agents for customer support, sales outreach, marketing, training simulations, and recruiting. * Gradium builds real-time speech models: TTS, STT, voice cloning, and voice design. Streaming-native, built for voice agents, with latency low enough for natural turn-taking and reliable pronunciation on alphanumerics. The voice library is tuned for agent use, with cloning and voice design for teams that need their own accents, pitches, and tones. Spun out of the Kyutai research lab in 2025, from the team behind Moshi. * Inception is the creator of Mercury 2, a new class of diffusion LLMs that match the quality of frontier speed-optimized models, with sub-250 ms time to first token, 1,000+ tokens per second throughput, and 70% lower cost per task. Served on Baseten, top use cases include real-time voice applications, search and retrieval pipelines, and AI coding sub-agents (e.g., context compaction). * NVIDIA provides open models such as Nemotron ASR and BioNeMo under permissive licenses for commercial use. For customers who want production-ready deployments, NVIDIA NIM packages optimized inference runtimes, standard APIs, and enterprise-grade operational capabilities. Through its distribution agreement with NVIDIA, Baseten makes it easy to deploy NIM on its inference platform while also respecting NVIDIA's commercial license. * PyannoteAI is the creator of Pyannote, a speaker intelligence foundation model for audio diarization. It can identify who spoke, when, and how, even in complex audio, and can pair with any STT model to turn transcripts into clean, speaker-attributed metadata. With centisecond boundary precision, Pyannote delivers 300ms real-time diarization latency, whether in cloud, on-premise, or at the edge, and up to 52% lower transcription errors. Top use cases include meeting transcription, healthcare AI scribing, financial compliance, and media dubbing. * SID.ai is the creator of SID-1, a specialized search LLM that excels at multi-step agentic retrieval across unstructured datasets. SID-1 surfaces ~2x more relevant documents than embedding search without requiring reindexing. Providing agentic search at ~100x lower cost and ~20x lower latency than frontier models, it powers high-recall retrieval for legal, healthcare, customer support, general knowledge, and more. * Subconscious builds specialized inference systems for long-horizon AI agents. Using near-lossless message compression and efficient prefix/suffix caching, it expands stateful context windows by 10x, increases throughput by 3.5x, and reduces token consumption by up to 50% without sacrificing accuracy. Serving its TIM models alongside open-weight options, top use cases include coding, workflow, and browser automation, research, and legal document generation. * Synthefy's Nori is a tabular foundation model that outperforms XGBoost in accuracy without requiring retraining, feature engineering, or hyperparameter tuning. Operating in seconds on a single GPU via Baseten through zero-shot in-context learning, its instant, in-context predictions streamline high-volume workflows like demand forecasting, fraud detection, predictive maintenance, risk scoring, churn protection, and high-frequency trading. The list is long already, but there are many more model labs that Baseten is proud to partner with: Bria, Canopy Labs (Orpheus TTS), Krea (Turbo), MongoDB (Voyage 4), Musubi (PolicyLM-1b), Scaled Cognition (APT-1). Baseten look forward to working with many more and seeing how its customers leverage these models to build awesome applications. The Baseten distribution platform is bringing together its lab partners and its customers. Baseten is excited to be at the center of this thriving ecosystem. The future is a model ecosystem. Baseten believe the future of AI won't run on just a few large frontier models. It will be powered by a diverse ecosystem of models, open and closed. Baseten is investing in a future where open-weight models coexist with dozens of specialized closed-weight models, each optimized for a specific domain, modality, or use case. For that future to thrive, the labs building those models need an efficient way to bring their innovation to production, reach the right customers, and deploy them securely at scale. Baseten for Model Labs makes this future possible. By combining production-ready inference infrastructure with distribution through the Baseten Model Library and joint go-to-market support, Baseten help model labs focus on building great models while Baseten power their inference growth for the long term. Talk to Baseten Connect with its product experts to see how Baseten can help.

Jenuel Dev
Jun 16th, 2026
Voice AI is becoming a full-stack problem.

Voice AI is becoming a full-stack problem. Cartesia's Sonic-3.5 and Ink-2 launch shows why voice AI is becoming a full-stack engineering problem, not just a nicer text-to-speech demo. #Speech Models #AI Engineering #Product Development Jun. 16, 2026. 1:04 AM The next useful AI app may not look like a chatbot at all. It may sound like a calm support rep, a patient tutor, or a field assistant that can listen while someone is busy with both hands. That is why Cartesia's new Sonic-3.5 and Ink-2 launch is worth paying attention to. The headline is not just another text-to-speech upgrade. The more important signal is that voice AI is turning into a full-stack engineering problem: speech-to-text, text-to-speech, latency, turn-taking, interruptions, safety checks, and tool calls all have to feel like one product. For builders, this is the shift. A voice agent is not a language model with audio bolted on. It is a real-time system where every extra delay makes the product feel less intelligent. What changed. Cartesia announced Sonic-3.5 for text-to-speech and Ink-2 for speech-to-text, positioning them as a paired stack for real-time voice agents. The company says Sonic-3.5 is built for naturalness, low latency, and 40+ languages, while Ink-2 focuses on transcription accuracy and fast turn-taking. The most practical claim is the pipeline framing. Cartesia is selling STT and TTS as parts of the same real-time loop instead of two separate vendors that developers have to stitch together. Its launch page points to sub-90ms TTS and 100ms transcript latency with native turn detection. If that holds up in real applications, it matters more than a demo voice sounding slightly nicer. Why developers should care. Voice agents fail in small moments. A half-second pause after every sentence feels robotic. Poor interruption handling makes users repeat themselves. Bad transcription turns a simple request into a support ticket. A beautiful generated voice is not enough if the agent cannot listen, stop, recover, and call tools quickly. This is where the developer opportunity is. The best voice products will not be built by choosing the most impressive model in isolation. They will be built by measuring the whole conversation loop. * Support agents: detect intent, pull order data, answer naturally, and escalate when confidence drops. * Healthcare and field workflows: capture spoken notes while the user is working, then structure them for review instead of pretending the transcript is final truth. * Education apps: let students talk through a problem and interrupt the tutor when they are confused. * Internal tools: create voice interfaces for dashboards, incident updates, and hands-free task capture. The common thread is not voice for novelty. It is voice where typing is slower, unsafe, or unnatural. The weak spots to test before shipping. I would not ship a production voice agent just because a model page says low latency. Builders should test the boring edge cases first. * Interruptions: can the user cut the agent off without the conversation state breaking? * Noisy audio: does transcription degrade gracefully in a cafe, car, warehouse, or cheap headset? * Accents and code-switching: does the system handle real users, not just studio samples? * Tool-call delay: what happens when the LLM and backend API take longer than the speech layer? * Consent and recording: is it clear when audio is captured, stored, or used for review? The mistake is treating speech as a UI skin. Voice changes the trust model. People reveal more when they speak, and they notice awkward timing faster than they notice a slow web page. A practical builder checklist. If you are evaluating Sonic-3.5, Ink-2, or any competing voice stack, build a small benchmark around your own product instead of relying on generic leaderboard claims. * Measure time from user speech ending to agent response starting. * Track word error rate on your real vocabulary, including names, product terms, and acronyms. * Test barge-in behavior: interrupt the agent mid-sentence and see if it adapts. * Log every failed turn with audio, transcript, intent, tool call, and final response. * Decide when the agent should stop talking and ask a human to take over. That last point is important. A good voice agent should not be endlessly confident. In many products, the trust-building moment is the handoff: 'I am not sure, so I am sending this to a person with the context attached.' The bigger signal. The AI industry has spent years making models that can answer. The next competition is around systems that can participate. Voice makes that obvious because participation has rhythm: listening, pausing, interrupting, confirming, and acting. Cartesia's launch is one signal in that direction. Whether its models become the default stack or not, the direction is clear: builders need to think less about isolated model calls and more about complete interaction loops. For developers, the question is no longer 'Can I add a voice mode?' The better question is: 'Where would a fast, interruptible, trustworthy conversation make this product meaningfully better?' References. Thanks for reading! If you enjoyed this article and like this kind of content, you're always welcome to buy me a little coffee, but only if you'd like to. No pressure at all, and either way I'm truly grateful you stopped by.

IANS
Mar 27th, 2026
Smallest.ai launches Lightning V3, a new text-to-speech model that beats OpenAI, Cartesia, and ElevenLabs on key voice quality benchmarks.

Smallest.ai launches Lightning V3, a new text-to-speech model that beats OpenAI, Cartesia, and ElevenLabs on key voice quality benchmarks. Vmpl. * March 27, 2026 6:55 AM Smallest.ai Designed for real-time use, it combines multilingual speech, voice cloning from seconds of audio, and conversational-level prosody in a single system San Francisco, CA | March 27, 2026 - Smallest.ai, the research-first Voice AI company building proprietary speech models and production-grade voice agents, today announced the launch of Lightning V3, its most advanced text-to-speech (TTS) model for real-time, conversational AI. In conversational evaluations, Lightning V3 achieves a 3.89 MOS, outperforming leading models from OpenAI, Cartesia, and ElevenLabs, while also leading on intonation (3.33) and prosody (3.07)- two of the most critical factors for natural, human-like speech. The model combines this performance with multilingual support, instant voice cloning, and streaming generation designed for real-world interactions. Most TTS models today are still evaluated on complete sentences generated in isolation. That setup is easier to optimize for, but it doesn't reflect how voice systems actually behave in production- where audio is generated in chunks, context is incomplete, and responses have to adapt as conversations unfold. Lightning V3 is built for how voice systems actually run in production- generating speech in chunks, without full context, and adapting as conversations evolve. It maintains consistency across turns and adjusts tone and pacing mid-sentence, which is where most systems break down. That same setup allows the model to work across use cases without retraining- including voice agents, contact centers, podcasts, audiobooks, dubbing, and interactive applications. It supports 15 languages with automatic detection and mid-sentence switching, and can clone a voice from 5-15 seconds of audio. These cloned voices tend to sound more natural than preset ones, since they retain the variations of real speech. The model outputs audio at 44.1 kHz, and can be downsampled to 8-24 kHz for telephony. "Conversation is where most voice systems fall apart," said Sudarshan Kamath, Founder and CEO, Smallest.ai. "It's not just about sounding clear- the voice has to track context, timing, and emotion at the same time. If it works there, it works everywhere." A shift in how voice quality is measured. The launch also challenges how voice models are evaluated. Most benchmarks rely on static outputs- a setup that rarely reflects real usage. Lightning V3 is evaluated across these use case specific settings, measuring how well the voice maintains coherence, responsiveness, and believability throughout an interaction, in the given context of the conversation not just within a single utterance. Voices should be designed and judged in context: for whether they fit the persona they are meant to inhabit, carry the right social signal, and feel believable in the moment they were built for. Pricing. Lightning V3.1 is available on a pay-as-you-go model, with no upfront commitments, seat licenses, or minimum usage requirements. Teams can scale from early prototypes to high-volume deployments across both voice agents and content generation- with usage-based pricing and non-expiring credits. About Smallest.ai. Smallest.ai is a research-first Voice AI company building proprietary speech models and production-grade voice agents for regulated enterprises. The company develops state of the art speech-to-text, text-to-speech, and real-time voice systems, enabling end-to-end automation of high-volume conversations across support, collections, onboarding, and servicing- without relying on stitched third-party APIs. Designed for financial services and other regulated industries, Smallest.ai is SOC 2, GDPR, HIPAA, and PCI compliant, supports on-prem and private cloud deployments, and operates reliably in multilingual environments. Its platform is used in production by enterprises across banking, insurance, BPO and telecommunications in the US and India. Disclaimer: The content provided in this section is part of a third party press release service and does not reflect the editorial views or opinions of IANS. The responsibility for the accuracy, authenticity, and legality of the information lies solely with the content provider. IANS assumes no liability for the content published under this arrangement and encourages readers to verify the information independently before consuming it.