Full-Time
Cloud-based ASR platform with multilingual transcription
No salary listed
London, UK
Hybrid
London hybrid role with 2-3 designated office days per week.
See people who can refer or advise you
Speechmatics provides automatic speech recognition (ASR) technology that converts spoken language into text. It offers a platform-as-a-service (PaaS) with an API, enabling developers and businesses to embed real-time or batch transcription into their apps and workflows. Deployment options include cloud, on-premises, and on-device to support different security needs. The models are trained on millions of hours of unlabeled audio to support many languages, dialects, and accents, including a Global English model that handles major English accents with a single backbone. Distinguishing features include speaker identification, automatic punctuation, translation, and products like Ursa and Flow. The company targets media, contact centers, enterprise communications, and healthcare, and uses a usage-based pricing model with tiers and custom enterprise licensing. Its goal is to make speech-to-text accessible and accurate for diverse voices across industries.
Company Size
51-200
Company Stage
Series B
Total Funding
$81.6M
Headquarters
Cambridge, United Kingdom
Founded
2009
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Dental Insurance
Flexible Work Hours
Hybrid Work Options
Paid Vacation
401(k) Company Match
401(k) Retirement Plan
Family Planning Benefits
Fertility Treatment Support
Home Office Stipend
Stenograph has partnered with Speechmatics to integrate speech recognition directly into CATalyst VP, its voice reporting software for court reporters. The collaboration addresses challenges voice reporters face managing multiple technology platforms. The integration eliminates the need for separate third-party speech recognition software, reducing issues like latency and system instability. Speechmatics runs entirely on-device, functioning without internet connectivity whilst keeping sensitive proceedings on the reporter's computer. "This exclusive partnership gives voice reporters the consistency, reliability, and ease of use they've been seeking within a single solution," said Michelle McLaughlin, Stenograph's Vice President of Sales. The integration removes requirements for additional devices and per-minute charges. Speechmatics will also power CheckIt Plus, an add-on service for stenographers and voice reporters to improve realtime transcription and reduce editing time.
Stenograph and Speechmatics announce industry-first on-device integration for CATalyst VP. The move brings speech recognition directly into CATalyst VP eliminating the challenges of running multiple applications, making it the first seamless solution for the voice reporting industry. Stenograph, the global leader in court reporting technology, today announced an exclusive integration with Speechmatics, bringing industry-leading on-device & realtime speech recognition directly into CATalyst VP, its specialized voice reporting software. As voice reporting continues to grow across the legal industry, professionals face the challenges of managing multiple technology platforms to produce an accurate record. Running separate speech recognition software alongside CAT software often requires additional logins, independent update schedules, compatibility checks, and ongoing troubleshooting - creating unnecessary complexity in a profession where focus and accuracy matter most. To address these challenges, Stenograph and Speechmatics have partnered to deliver the first fully integrated speech recognition solution built directly into Voice Reporting CAT software. This exclusive integration eliminates the need for voice reporters to run a separate third-party speech engine alongside CATalyst VP, creating a more streamlined and accurate user experience while helping reduce common issues such as latency, system instability, and software version conflicts. Because Speechmatics runs entirely on-device, that reliability holds even without an internet connection, so reporters get the same performance in a fully connected courtroom or an offline deposition, with sensitive proceedings never leaving a voice reporter's computer. For decades, court reporters have trusted Stenograph to provide the tools and services they need to take down and produce the verbatim record. This exclusive partnership with Speechmatics gives voice reporters the consistency, reliability, and ease of use they've been seeking - all within a single solution. By eliminating the need for separate speech recognition software, high-spec computers, additional devices, or an internet connection, CATalyst VP with Speechmatics allows voice reporters to focus more confidently on capturing the record. It also removes the uncertainty of per-minute charges. - Michelle McLaughlin, Vice President of Sales at Stenograph. California official court reporter and CATalyst VP user, Vance Malone, CVR-M, RVR, stated: What excites me most isn't just another speech engine - it's having Speechmatics fully integrated into CATalyst VP. Native integration removes one of the biggest historical challenges for voice reporters: getting multiple applications to work together reliably in realtime. Giving reporters another high-quality recognition option, while keeping Dragon available, provides flexibility without adding complexity. In addition to powering CATalyst VP for voice reporters, Speechmatics will also power CheckIt Plus, an add-on service used by both stenographers and voice reporters to deliver cleaner realtime and rough drafts while reducing overall editing time. Court reporters need technology that truly simplifies their work. Building its speech recognition directly into Stenograph's software means legal professionals get that accuracy, connectivity and privacy they need without adding the complexity of another tool to manage. - Ricardo Herreros-Symons, CRO of Speechmatics. " You can learn more at StenographxSpeechmatics PR Media contacts: Mieke Smith [email protected] Madi Dixon [email protected]
Introducing Melia, its new multilingual speech-to-text model. A multilingual speech-to-text model from Speechmatics, with code-switching across 56+ languages. Available today in production preview, starting with batch transcription. Tl;dr. * Melia is a multilingual speech-to-text model with code-switching across all 56+ languages Speechmatics Limited support, live today in production preview for batch. * It's strongest on accented, code-switched speech, and beats Deepgram, Microsoft, and AssemblyAI on most FLEURS languages. * It's its lowest-priced model (from $0.129/hour, 10 hours/month free) and sits alongside Standard and Enhanced rather than replacing them. Today Speechmatics Limited is launching Melia, a multilingual model that handles code-switching across all 56+ languages Speechmatics Limited support. It's in your Speechmatics account now, in the Portal, Batch API, and SDKs. Speechmatics has always been driven by one mission: to understand every voice. For over two decades, that has meant the most accurate real-world speech-to-text available, especially for the accents and dialects other systems treat as edge cases. But until now, it has largely meant one language at a time. Over half the world speaks more than one language, and people move between them mid-conversation. For a growing number of its customers, understanding every voice now means transcribing several languages in a single file, quickly and affordably. Melia is where that begins, in batch today, with more to come. What Melia does differently and where Speechmatics Limited is heading Melia covers all the languages Speechmatics Limited support, so a recording that moves between languages comes back as one continuous transcript, with no language to select in advance and no language packs to manage. That keeps your workflow, and the orchestration behind it, simple. Handling multiple languages is one thing. Handling how multilingual people actually sound is another: they carry an accent from one language into the next, and that accented speech is where most models struggle. It's where Speechmatics' two decades of accuracy work shows up, and where Melia is strongest. Its goal for this lineage is to build the world's most accurate code-switching speech-to-text model. Melia 1 is the first step toward that, and each release from here makes its code-switching more accurate, with further improvements landing regularly over the coming weeks. How it compares with other providers On FLEURS, an open benchmark across many languages of read speech, Melia produces fewer errors than Deepgram, Microsoft and AssemblyAI on most languages. Measured against each vendor's strongest model, here's the share of languages where Melia wins: * Deepgram: 91% (best of Nova-3 and Nova-2) * Microsoft: 91% (best of Enhanced and Standard) * AssemblyAI: 77% (best of Universal-3-Pro, Universal-2, and Universal) Where Melia fits with its current models Melia sits alongside Standard and Enhanced, not in place of them: it's the multilingual addition to the lineup. When the lowest possible word error rate is what matters, Enhanced is still the model to choose. For multilingual audio, Melia is the obvious option. It's its lowest-priced model, and turnaround is blisteringly fast and getting faster. But it's not only for multilingual work. In its internal benchmarks on challenging, noisy monolingual audio, Melia averages a 5% lower word error rate than Standard across the languages Speechmatics Limited tested. For many single-language workloads, including ones running on Standard today, it's worth testing Melia in its place, at a lower price. Standard and Enhanced still offer more features, including real-time transcription. See the documentation to compare all three. Ready for production workloads: Melia is part of the same Batch API as Standard and Enhanced, and runs on the same production infrastructure. You get the same cloud regions, plus on-prem for teams that run Speechmatics in their own environment. The SLAs and reliability you depend on apply to Melia from launch. And because it covers every language in one model, there's just one to integrate, manage, and run instead of one per language. Melia carries a preview label because it's improving quickly, with more to come over the coming weeks and months, and Speechmatics Limited welcome feedback from production users. The three examples below are a starting point: places where Melia already makes a real difference today, and far from the only ones. Contact center analytics: Multilingual call recordings that monolingual models couldn't process are now transcribable at scale, across European, Gulf, Southeast Asian, and US Hispanic operations. Melia returns language metadata with every transcript, so you can break calls down by language for routing, reporting, and quality work. Multilingual broadcast captioning: Spanish-English content for US Latino audiences, Arabic-English broadcasts across the Gulf, multi-language news from across Southeast Asia: one model, not one per language. Compliance monitoring: Regulated teams record customer calls across dozens of markets and have to review all of them. Melia transcribes the full archive, including the language switches that leave monolingual transcripts with gaps, so reviews aren't built on partial records. It also labels the languages in each call, so every one can reach a reviewer who speaks them. What's next? Melia is early in its life and will keep improving quickly. Most of that work is on accuracy: building on the gains you see here, making code-switching sharper, and expanding features and functionality. Real-time is on the way too, first in preview and then in production, so the same model can handle live audio as well as files. Expect regular improvements rather than big releases. Team members on Melia's impact, use cases, vision, and what's coming. Try it Melia is live in your Speechmatics account today. Select Melia 1 in the Portal, or set melia-1 in your Batch API config. It's its lowest-priced model: as low as $0.129 per hour, with 10 hours per month free. Volume and enterprise pricing available; talk to your account team. Documentation, supported languages, and SDK code examples: docs.speechmatics.com.
Adobe and Speechmatics deliver cloud-grade speech recognition on-device for Premiere. GlobeNewswire | Speechmatics Limited Today at 6:07am PDT CAMBRIDGE, United Kingdom, April 21, 2026 (GLOBE NEWSWIRE) - Dialogue is the centerpiece of modern content. Whether it's a podcast, a DIY instructional video, or a documentary, what people say drives the story. Accurately understanding speech and giving creators control over how it's used has become essential to producing compelling, high-quality content. Now, as LLM-centric workflows take hold and natural language becomes the interface for shaping stories, that speech-to-text foundation matters more than ever. Accurate transcription isn't just a feature - it's the layer that optimises content workflows, enables faster content creation and makes agentic AI work. Speechmatics has been Adobe's partner since 2021, when Adobe became the first non-linear editing platform to include speech-to-text (STT) in Premiere. Today, that partnership deepens with a new on-device STT model in Premiere that delivers near-cloud accuracy while keeping all audio local to the device. ON-DEVICE FROM THE START, EVOLVED FOR TODAY When Adobe launched STT for Premiere, large enterprises couldn't always use cloud-based services due to privacy concerns. Speechmatics was one of the few providers with on-device models - a key reason for the partnership. Five years later, those privacy requirements haven't changed. With the rise of LLMs and data sovereignty concerns, the need for secure deployments has, in fact, increased. What has changed is the performance gap: Speechmatics' new on-device model brings local transcription on par with cloud accuracy with optimisations to run efficiently. Studios, agencies, and production companies handling content before it goes public can now work seamlessly from anywhere: on a film set, between client meetings, on a flight - at full accuracy, with no dependency on a connection and no interruption to the work. Editing video and audio with text, creating captions quickly, and labeling speakers with industry-leading speaker diarization - all local, all private, all accurate. VOICE AI THAT WORKS FOR EVERYONE For voice to be useful for creative work, it has to understand how people actually speak. The new Speechmatics on-device model has been trained on millions of hours of speech to deliver high accuracy for accented speech, non-native speakers, and noisy environments like field reporting or film sets. The benchmark results reflect that. The new on-device model in Premiere: * Is within 5% relative to cloud accuracy, evaluated across nearly 10 million words of diverse real-world data * Processes 1 hour of audio in about 55 seconds * Leads the way against the closest competitor, with a 12-16% improvement against Whisper-powered creative solutions * Runs on Windows & Mac, making use of the latest AI acceleration techniques to ensure efficient processing across a range of hardware, including broad hardware support for the latest Mac M5, NVIDIA RTX, AMD GPUs and older hardware such as Intel Macs "Adobe's global creator community speaks hundreds of languages and dialects. Since 2021, our partnership has focused on making sure speech technology works for everyone - whether you're editing in Scottish English, Mexican Spanish, or Cantonese. Today, millions of users can benefit from accurate transcription that works anywhere - on-device for privacy, and in the cloud for scale - without compromising performance. As Adobe builds toward LLM-powered creative workflows, having a speech foundation that truly understands diverse voices becomes even more critical. We're proud to be part of that future." - Katy Wigdahl, CEO, Speechmatics. AVAILABILITY Speechmatics on-device joins Speechmatics cloud and Speechmatics on-prem as a purpose-built option for ISVs and OEMs where data residency, offline capability, or predictable costs make local execution the right architectural call. It integrates as a C/C++ library on macOS and Windows. ABOUT SPEECHMATICS: Speechmatics is the Voice AI company on a mission to understand every voice. Its speech-to-text technology delivers industry-leading accuracy across 55+ languages, with specialized models for healthcare, media, contact centers, and enterprise organizations worldwide. Speechmatics powers leading technology providers including Adobe, AI Media, Content Guru, and Nordhealth, and offers deployment across cloud, Speechmatics on-prem, and on-device. Headquartered in Cambridge, UK. Learn more at www.speechmatics.com. MEDIA CONTACT: Mieke Smith // Content Lead, Speechmatics // [email protected] This is a paid placement. For further inquiries, please contact GlobeNewswire directly.
Speechmatics and thymia have partnered to combine medical-grade speech-to-text with clinical-grade voice biomarker intelligence, identifying over 30 health signals from just 15 seconds of natural speech. The platform detects stress, fatigue, depression, anxiety symptoms, type 2 diabetes and driver impairment at over 85% accuracy through a single integration. thymia's technology, built on a dataset of 75,000-plus unique voices and validated in over 20 peer-reviewed publications including Nature Scientific Reports, processes audio at 15-second intervals. Speechmatics' speech engine ensures accuracy across accented speech and non-native speakers. The technology is already deployed in safety-critical industries including automotive. The voice biomarker market is projected to reach $5 billion by 2028, whilst the EU mandates driver fatigue monitoring in all new vehicles by July 2026.