Full-Time

Member of Technical Staff

Agent Platform, Agent OS

Boson Ai

Boson Ai

11-50 employees

Develops scalable AI tools for enterprises

Compensation Overview

$150k - $400k/yr

Santa Clara, CA, USA

In Person

Category
Software Engineering (1)
Required Skills
React.js
RAG
Observability
LangChain

Get referred to Boson Ai

See people who can refer or advise you

Requirements
  • Deep Experience: 3+ years of hands-on experience in backend engineering and distributed systems, with a track record of building and owning core platforms or frameworks used successfully by other engineering teams.
  • Agentic Systems Expertise: Demonstrated, hands-on experience architecting, building, or operating production-grade agentic systems: orchestrating LLM calls, managing complex tool interactions, and defining stateful workflows—moving beyond simple single prompt/response API integrations.
  • Orchestration & Design Patterns: Strong working knowledge of engineering orchestration frameworks (e.g., LangChain, LlamaIndex, or internal equivalents) and a deep understanding of core design patterns like RAG, ReAct, and multi-step planning.
  • Systems Engineering Mastery: Deep and practical understanding of distributed system design, concurrent programming, and building for reliability in multi-tenant cloud environments with strictly defined latency and cost envelopes.
  • Framework Evangelism: Proven experience designing, implementing, and rolling out successful frameworks or libraries that other internal engineering teams enthusiastically adopt and productively build upon.
  • Security Focus: Comfort and prior experience working on security-sensitive systems, including implementing authz/authn schemes, isolation boundaries, data protection protocols, and integrating with centralized policy/safety infrastructure.
  • Technical Leadership: Strong technical communication skills and the ability to lead complex, cross-functional technical initiatives, driving consensus and influencing architectural decisions across partner teams.
Responsibilities
  • System Ownership: Take ownership of the core dialog & policy engine. Define and implement the state machine for agent state representation, the decision-making logic, and the mechanisms for enforcing complex safety policies and guardrails at the execution layer of a workflow.
  • Distributed Context & Memory: Design, implement, and maintain the high-performance context and memory systems. Focus on low-latency, reliable access to conversational and user history, including the tight integration and optimization of RAG and vector retrieval pipelines for production use.
  • Agentic Orchestration Frameworks: Define, architect, and deliver robust agentic orchestration patterns, including battle-tested planner–executor schemes, ReAct-style reasoning and acting loops, and resilient, multi-step workflows that programmatically combine tools, LLMs, and stateful memory.
  • Internal SDK/Framework Development: Build and evolve the internal, production-grade equivalent of frameworks like LangChain/LlamaIndex. Design composable graphs and execution chains with clear APIs and type safety that product engineering teams and low-code builders can safely reuse, extend, and deploy at scale.
  • Voice Runtime Infrastructure: Own and optimize the voice runtime components for streaming audio, low-latency barge-in detection, and reliable turn-taking protocols. This requires deep collaboration with Application and ML Platform teams to meet tight latency, jitter, and quality of service (QoS) constraints.
  • Tooling & Integration Architecture: Architect a robust, secure tooling and integration framework (MCP/A2A). This includes building the underlying infrastructure for tool registration, handling complex authentication/authorization, implementing rate limiting/circuit breaking, managing retries, and ensuring typed, validated I/O between agents and external microservices.
  • Platform Observability & Reliability: Define, instrument, and monitor rigorous SLIs/SLOs for the Agent Platform. Lead engineering efforts to continuously improve reliability, enhance system debuggability (rich, step-level traces and structured logging), and drive core performance optimizations over time.
  • API & Abstraction Design: Ensure the platform's public-facing APIs and internal abstractions are clear, well-documented, and fundamentally sound, enabling junior and senior engineers alike to compose sophisticated agent behavior without introducing systemic invariants or breaking changes.
  • Advanced Capabilities R&D: Explore and prototype future capabilities, focusing on the engineering challenges of on-device personalization, implementing privacy-preserving federated learning signals, or integrating novel policy adaptation techniques that influence agent behavior in production.
Desired Qualifications
  • Experience developing and operating conversational AI platforms, agent frameworks, or high-throughput, complex workflow engines in a production setting.
  • Engineering background in real-time media (audio/video) systems or low-level signaling protocols where extreme low-latency and jitter management are critical performance factors.
  • Prior experience building high-stakes enterprise platforms (e.g., payments, identity, core data services) where correctness, auditability, and absolute reliability are non-negotiable requirements.
  • Exposure to emerging systems and engineering techniques, such as integrating federated learning models, enabling on-device personalization, or implementing bandit-style adaptive policy systems.

Boson AI develops large language model tools to power AI-driven experiences in virtual worlds. Its products understand and generate human-like text and are designed for wide use, from individuals to large enterprises. The tools work by combining advanced deep learning with system engineering to create customizable LLM-based applications that can be embedded into various software to personalize storytelling, learning, content creation, and data insights. Boson AI differentiates itself by focusing on tailored experiences in virtual environments across multiple industries, offering scalable solutions through product sales, subscriptions, and licensing. The company’s goal is to provide practical, personalized AI tools that enhance user interactions, storytelling, education, and business intelligence in virtual settings, becoming a leading provider in the AI market.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

Santa Clara, California

Founded

2023

Get referred to Boson Ai

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Boson AI raised $70 million, giving runway for model training and enterprise sales.
  • Fortune reported August 2026 enterprise targeting in finance, telecom, healthcare, and insurance.
  • Higgs TTS 3 and Higgs Avatar launched in June 2026, showing rapid product expansion.

What critics are saying

  • OpenAI, Meta, and Google can bundle voice into existing platforms and crush Boson distribution.
  • Higgs Realtime's beta timing and unfinished docs risk adoption delays before enterprise renewals in 2026.
  • If Boson cannot sustain model-quality leadership, its low-price pitch becomes a commodity trap.

What makes Boson Ai unique

  • Alex Smola launched Higgs Realtime on August 21, 2026, for live speech-to-speech agents.
  • Boson AI claims one-tenth competitor costs, with OpenAI-compatible realtime APIs and 100+ languages.
  • The company combines TTS, ASR, and avatar models into one voice-and-visual agent stack.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Growth & Insights and Company News

Headcount

6 month growth

8%

1 year growth

2%

2 year growth

22%
Boson AI
Aug 20th, 2026
Building voice AI that feels live: a hands-on guide to Higgs Realtime.

Building voice AI that feels live: a hands-on guide to Higgs Realtime. The Boson AI Team August 20, 2026 Real-time voice AI is easy to demo. Building it well is much harder. A good voice agent can't simply wait for someone to finish speaking, turn the audio into text, generate an answer, and read it back. Real conversations don't behave like a sequence of clean requests and responses. People pause. They interrupt. They change direction halfway through a sentence. They expect the system to remember what came before, use tools when necessary, and respond without making the conversation feel like a series of transactions. That changes how developers need to think about building voice applications - so today Boson is releasing the Higgs Realtime API Tutorial, a build-it-yourself guide for developers who want to understand how real-time voice AI works and build voice applications with sub-second latency at a fraction of the usual cost. Rather than starting with a finished demo and hiding the complexity underneath it, the tutorial builds a working browser voice assistant one layer at a time. From API calls to a live conversation. Traditional AI applications have a familiar rhythm: request, inference, response. Real-time voice runs on a different model of the world.

ScitiX
Jul 9th, 2026
Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference.

Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference. Venus Yang Speech is the most natural interface there is - and it's quickly becoming the default way people interact with AI. Today ScitiX is excited to announce that two of Boson AI's flagship speech models are available on the ScitiX Model Inference platform: bosonai/tts for text-to-speech, and bosonai/asr for speech recognition. Together they give you both halves of a voice experience - listening and speaking - behind a single API, on infrastructure that's built to be simple, fast, and secure. Meet the Boson AI models. Boson TTS is built for voice chat. It doesn't just read text aloud - it speaks, producing expressive, conversational speech that sounds like a person rather than a narrator. * Expressive & conversational - natural emotion, style, and prosody, with inline control * 100+ languages out of the box * Zero-shot voice cloning - capture a voice from a short sample, no fine-tuning required * $0.05 / minute If you're building voice assistants, agents, audiobooks, or any product where tone matters, bosonai/tts is designed to make the output feel alive. Boson ASR is a state-of-the-art speech recognition model that turns spoken audio into accurate text - reliably, across languages, and even when conditions aren't ideal. * State-of-the-art accuracy on real-world audio * 90+ languages supported * Robust in noisy environments - call centers, mobile, the real world * Streaming support for low-latency, real-time transcription * $0.006 / minute Pair the two and you have a full speech loop: bosonai/asr to understand what users say, bosonai/tts to respond in a natural voice. A platform designed to get you live in seconds. Great models are only useful if you can actually ship them. That's the whole point of ScitiX Model Inference - and it shows up in three places. Simple - discover, view, integrate. Every model on the platform lives in the Model Plaza, a searchable marketplace you can filter by provider, type, context window, and price. The flow is the same for every model: * Find the model - e.g. the bosonai/tts card in the Plaza. * View Details - read the spec, capabilities, pricing, and a ready-to-run code example. * Copy the endpoint and key into your config, set the model name, and you're done. js// config.js - model integration export default {baseURL: 'https://api.scitix.ai/model-api', apiKey: 'sk-scitix- - - - - - - - - - - - - ', model: 'bosonai/tts', // swap to 'bosonai/asr' for speech recognition} The API is OpenAI-compatible, so it slots into the tools and SDKs you already use. Fast - production-ready in seconds, not weeks. There's no provisioning, no model download, no infrastructure to stand up. From browsing the catalog to your first API call is a matter of seconds - any model in the Plaza is production-ready the moment you copy its endpoint and key. Behind the scenes, ScitiX handles the GPU capacity and scaling so your latency stays low as your traffic grows. And switching between models is trivial: the endpoint and authentication never change - the only thing that varies between bosonai/tts and bosonai/asr is the model name in your request. In its internal benchmarks, Boson TTS stays responsive across hardware tiers - first audio comes back in tens of milliseconds, with steady streaming throughput: Note: these figures are from internal test runs (Boson TTS streaming, single concurrent request, 100-sample suite) and are provided for reference only. They are test data, not a performance commitment or SLA, and may vary by workload, configuration, and region. Secure - your keys, your control. Authentication is a single Authorization: Bearer header carrying your API key. On the platform side, keys are always masked, can be rotated or revoked at any time, and account access can be protected with an extra layer of two-factor security plus security notifications. Your credentials stay yours. On the infrastructure side, your workloads run on ScitiX's own footprint: ScitiX operates Tier III+ (T3+) data centers in North America, and expects to complete its SOC 2 Type II certification by the end of 2026 - so the security story extends from your API key all the way down to the facilities your models run in. The takeaway. Boson AI's TTS and ASR give you a complete voice loop - expressive, natural speech and accurate, multilingual transcription - behind one OpenAI-compatible API. ScitiX Model Inference makes that loop simple to adopt, fast to scale, and secure by default, from your API key all the way down to the data center. Whether you're giving a product a natural-sounding voice or transcribing speech across 90+ languages, both models are live on ScitiX today - ready the moment you are.