Full-Time

Software Engineer

Platform & Application

Boson Ai

Boson Ai

11-50 employees

Develops scalable AI tools for enterprises

Compensation Overview

$150k - $270k/yr

Santa Clara, CA, USA

In Person

On-site at Santa Clara, California.

Category
Software Engineering (1)
Required Skills
Kubernetes
Rust
Python
Git
Java
ETL
Docker
AWS
Go
Observability
REST APIs
C/C++
DevOps
Google Cloud Platform

Get referred to Boson Ai

See people who can refer or advise you

Requirements
  • 0–2 years of professional software engineering experience (internships, co-ops, and strong personal or open-source projects count), or a recent CS degree with equivalent hands-on work.
  • Solid programming fundamentals and a genuine interest in backend and distributed systems — you understand concepts like concurrency and fault tolerance and are eager to apply them in production.
  • Some exposure to building backend services, APIs, or data processing — through work, coursework, or projects.
  • Proficiency in at least one language such as Python, Go, Java, Rust, or C++, and a willingness to pick up new ones.
  • Familiarity with the basics of cloud infrastructure (AWS/GCP), containers (Docker/K8s), version control, and CI/CD — or clear enthusiasm to learn them quickly.
  • Curiosity, strong communication, and a collaborative mindset — you ask good questions, welcome feedback, and want to grow.
Responsibilities
  • Contribute to the core platform infrastructure: the API serving layer, state management, policy enforcement, and execution runtime for agentic workflows — starting with well-scoped components and taking on broader ownership as you grow.
  • Help build and maintain distributed services that back our model API products, including request routing, rate limiting, and multi-tenant isolation, under the guidance of senior engineers.
  • Build and maintain pieces of our data pipelines (ETL/ELT) for API logs, usage analytics, and billing, with a focus on data correctness and freshness.
  • Develop and improve internal SDKs and libraries — writing clean, well-tested code with clear contracts that product teams can rely on.
  • Support our context and memory systems for conversational workloads: retrieval, caching, and integration with vector stores and retrieval pipelines.
  • Add observability across the platform — structured logging, tracing, and metrics — and help investigate and resolve reliability issues.
  • Collaborate with ML and product teams to integrate model serving, voice runtime, and tooling infrastructure, learning how the full stack fits together.
Desired Qualifications
  • Coursework, projects, or internship experience touching LLM serving, retrieval, or agentic systems.
  • Exposure to data pipeline tools (Kafka/Kinesis, Spark/Flink, Airflow) or agent frameworks.
  • Experience with real-time media (audio/video streaming) or any latency-sensitive system.
  • A track record of shipping something end-to-end — a side project, a hackathon build, or an open-source contribution you're proud of.

Boson AI develops large language model tools to power AI-driven experiences in virtual worlds. Its products understand and generate human-like text and are designed for wide use, from individuals to large enterprises. The tools work by combining advanced deep learning with system engineering to create customizable LLM-based applications that can be embedded into various software to personalize storytelling, learning, content creation, and data insights. Boson AI differentiates itself by focusing on tailored experiences in virtual environments across multiple industries, offering scalable solutions through product sales, subscriptions, and licensing. The company’s goal is to provide practical, personalized AI tools that enhance user interactions, storytelling, education, and business intelligence in virtual settings, becoming a leading provider in the AI market.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

Santa Clara, California

Founded

2023

Get referred to Boson Ai

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Higgs STT 3 reached #2 on OpenASR leaderboard, beating whisper-v3-large across 94 languages [2][5].
  • Higgs Avatar API launched June 2026 enables talking-head video generation from images plus audio or text [4][5].
  • Secured $70M in May 2025 seed funding to scale AI communication infrastructure and enterprise agent development [8].

What critics are saying

  • GPT-4o-transcribe and ElevenLabs Whisper-large dominate speech-to-text benchmarks, eroding Boson ASR adoption within 6–12 months [4].
  • Hume AI captures enterprise clients with superior chain-of-thought audio analytics and multi-turn dialogue fidelity in 9–15 months [3][5].
  • OpenAI's GPT-5o-audio will integrate native TTS/ASR with zero-shot cloning, undercutting Boson's $0.05/minute pricing in 12–18 months [2][5].

What makes Boson Ai unique

  • Boson AI focuses on personalized experiences in virtual worlds using deep learning and system engineering [3].
  • The company offers full-stack Higgs multimodal audio models plus Feynman Flow agentic platform for deployed intelligence [4].
  • Founded by AWS leaders Dr. Alex Smola and Dr. Mu Li, leveraging their Dive into Deep Learning expertise [2].

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Growth & Insights and Company News

Headcount

6 month growth

11%

1 year growth

0%

2 year growth

18%
ScitiX
Jul 9th, 2026
Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference.

Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference. Venus Yang Speech is the most natural interface there is - and it's quickly becoming the default way people interact with AI. Today ScitiX is excited to announce that two of Boson AI's flagship speech models are available on the ScitiX Model Inference platform: bosonai/tts for text-to-speech, and bosonai/asr for speech recognition. Together they give you both halves of a voice experience - listening and speaking - behind a single API, on infrastructure that's built to be simple, fast, and secure. Meet the Boson AI models. Boson TTS is built for voice chat. It doesn't just read text aloud - it speaks, producing expressive, conversational speech that sounds like a person rather than a narrator. * Expressive & conversational - natural emotion, style, and prosody, with inline control * 100+ languages out of the box * Zero-shot voice cloning - capture a voice from a short sample, no fine-tuning required * $0.05 / minute If you're building voice assistants, agents, audiobooks, or any product where tone matters, bosonai/tts is designed to make the output feel alive. Boson ASR is a state-of-the-art speech recognition model that turns spoken audio into accurate text - reliably, across languages, and even when conditions aren't ideal. * State-of-the-art accuracy on real-world audio * 90+ languages supported * Robust in noisy environments - call centers, mobile, the real world * Streaming support for low-latency, real-time transcription * $0.006 / minute Pair the two and you have a full speech loop: bosonai/asr to understand what users say, bosonai/tts to respond in a natural voice. A platform designed to get you live in seconds. Great models are only useful if you can actually ship them. That's the whole point of ScitiX Model Inference - and it shows up in three places. Simple - discover, view, integrate. Every model on the platform lives in the Model Plaza, a searchable marketplace you can filter by provider, type, context window, and price. The flow is the same for every model: * Find the model - e.g. the bosonai/tts card in the Plaza. * View Details - read the spec, capabilities, pricing, and a ready-to-run code example. * Copy the endpoint and key into your config, set the model name, and you're done. js// config.js - model integration export default {baseURL: 'https://api.scitix.ai/model-api', apiKey: 'sk-scitix- - - - - - - - - - - - - ', model: 'bosonai/tts', // swap to 'bosonai/asr' for speech recognition} The API is OpenAI-compatible, so it slots into the tools and SDKs you already use. Fast - production-ready in seconds, not weeks. There's no provisioning, no model download, no infrastructure to stand up. From browsing the catalog to your first API call is a matter of seconds - any model in the Plaza is production-ready the moment you copy its endpoint and key. Behind the scenes, ScitiX handles the GPU capacity and scaling so your latency stays low as your traffic grows. And switching between models is trivial: the endpoint and authentication never change - the only thing that varies between bosonai/tts and bosonai/asr is the model name in your request. In its internal benchmarks, Boson TTS stays responsive across hardware tiers - first audio comes back in tens of milliseconds, with steady streaming throughput: Note: these figures are from internal test runs (Boson TTS streaming, single concurrent request, 100-sample suite) and are provided for reference only. They are test data, not a performance commitment or SLA, and may vary by workload, configuration, and region. Secure - your keys, your control. Authentication is a single Authorization: Bearer header carrying your API key. On the platform side, keys are always masked, can be rotated or revoked at any time, and account access can be protected with an extra layer of two-factor security plus security notifications. Your credentials stay yours. On the infrastructure side, your workloads run on ScitiX's own footprint: ScitiX operates Tier III+ (T3+) data centers in North America, and expects to complete its SOC 2 Type II certification by the end of 2026 - so the security story extends from your API key all the way down to the facilities your models run in. The takeaway. Boson AI's TTS and ASR give you a complete voice loop - expressive, natural speech and accurate, multilingual transcription - behind one OpenAI-compatible API. ScitiX Model Inference makes that loop simple to adopt, fast to scale, and secure by default, from your API key all the way down to the data center. Whether you're giving a product a natural-sounding voice or transcribing speech across 90+ languages, both models are live on ScitiX today - ready the moment you are.