Full-Time

Site Reliability Engineer

Boson Ai

Boson Ai

11-50 employees

Develops scalable AI tools for enterprises

Compensation Overview

CA$125k - CA$250k/yr

Toronto, ON, Canada

Remote

Category
DevOps & Infrastructure (1)
Required Skills
Kubernetes
CUDA
Linux/Unix

Get referred to Boson Ai

See people who can refer or advise you

Requirements
  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
  • Strong hands-on expertise in cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
  • Strong hands-on expertise in distributed storage, particularly Ceph
  • Strong hands-on expertise in GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
  • Strong hands-on expertise in AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
Responsibilities
  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Desired Qualifications
  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure

Boson AI develops large language model tools to power AI-driven experiences in virtual worlds. Its products understand and generate human-like text and are designed for wide use, from individuals to large enterprises. The tools work by combining advanced deep learning with system engineering to create customizable LLM-based applications that can be embedded into various software to personalize storytelling, learning, content creation, and data insights. Boson AI differentiates itself by focusing on tailored experiences in virtual environments across multiple industries, offering scalable solutions through product sales, subscriptions, and licensing. The company’s goal is to provide practical, personalized AI tools that enhance user interactions, storytelling, education, and business intelligence in virtual settings, becoming a leading provider in the AI market.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

Santa Clara, California

Founded

2023

Get referred to Boson Ai

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Boson raised $70 million across two seed rounds, funding aggressive model development.
  • Higgs Audio 3.0 launched March 18, 2026, showing active product cadence.
  • Hackathons in Mountain View and Toronto signal developer traction and hiring pipeline.

What critics are saying

  • OpenAI, ElevenLabs, and Google dominate speech models, compressing Boson AI's pricing power.
  • Boson still lacks visible enterprise customer logos, signaling weak moat and fragile demand.
  • A product miss in 2026 would starve a seed-funded company before series A.

What makes Boson Ai unique

  • Alex Smola and Mu Li anchor Boson AI with deep-dive deep-learning credibility.
  • Higgs STT 3 supports 94 languages and outperforms Whisper v3 large on key languages.
  • Boson distributes through Eigen AI, ByteCompute, and ScitiX, broadening enterprise access.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Growth & Insights and Company News

Headcount

6 month growth

5%

1 year growth

2%

2 year growth

18%
ScitiX
Jul 9th, 2026
Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference.

Voice, built in: Boson AI's TTS and ASR models are now live on ScitiX Model Inference. Venus Yang Speech is the most natural interface there is - and it's quickly becoming the default way people interact with AI. Today ScitiX is excited to announce that two of Boson AI's flagship speech models are available on the ScitiX Model Inference platform: bosonai/tts for text-to-speech, and bosonai/asr for speech recognition. Together they give you both halves of a voice experience - listening and speaking - behind a single API, on infrastructure that's built to be simple, fast, and secure. Meet the Boson AI models. Boson TTS is built for voice chat. It doesn't just read text aloud - it speaks, producing expressive, conversational speech that sounds like a person rather than a narrator. * Expressive & conversational - natural emotion, style, and prosody, with inline control * 100+ languages out of the box * Zero-shot voice cloning - capture a voice from a short sample, no fine-tuning required * $0.05 / minute If you're building voice assistants, agents, audiobooks, or any product where tone matters, bosonai/tts is designed to make the output feel alive. Boson ASR is a state-of-the-art speech recognition model that turns spoken audio into accurate text - reliably, across languages, and even when conditions aren't ideal. * State-of-the-art accuracy on real-world audio * 90+ languages supported * Robust in noisy environments - call centers, mobile, the real world * Streaming support for low-latency, real-time transcription * $0.006 / minute Pair the two and you have a full speech loop: bosonai/asr to understand what users say, bosonai/tts to respond in a natural voice. A platform designed to get you live in seconds. Great models are only useful if you can actually ship them. That's the whole point of ScitiX Model Inference - and it shows up in three places. Simple - discover, view, integrate. Every model on the platform lives in the Model Plaza, a searchable marketplace you can filter by provider, type, context window, and price. The flow is the same for every model: * Find the model - e.g. the bosonai/tts card in the Plaza. * View Details - read the spec, capabilities, pricing, and a ready-to-run code example. * Copy the endpoint and key into your config, set the model name, and you're done. js// config.js - model integration export default {baseURL: 'https://api.scitix.ai/model-api', apiKey: 'sk-scitix- - - - - - - - - - - - - ', model: 'bosonai/tts', // swap to 'bosonai/asr' for speech recognition} The API is OpenAI-compatible, so it slots into the tools and SDKs you already use. Fast - production-ready in seconds, not weeks. There's no provisioning, no model download, no infrastructure to stand up. From browsing the catalog to your first API call is a matter of seconds - any model in the Plaza is production-ready the moment you copy its endpoint and key. Behind the scenes, ScitiX handles the GPU capacity and scaling so your latency stays low as your traffic grows. And switching between models is trivial: the endpoint and authentication never change - the only thing that varies between bosonai/tts and bosonai/asr is the model name in your request. In its internal benchmarks, Boson TTS stays responsive across hardware tiers - first audio comes back in tens of milliseconds, with steady streaming throughput: Note: these figures are from internal test runs (Boson TTS streaming, single concurrent request, 100-sample suite) and are provided for reference only. They are test data, not a performance commitment or SLA, and may vary by workload, configuration, and region. Secure - your keys, your control. Authentication is a single Authorization: Bearer header carrying your API key. On the platform side, keys are always masked, can be rotated or revoked at any time, and account access can be protected with an extra layer of two-factor security plus security notifications. Your credentials stay yours. On the infrastructure side, your workloads run on ScitiX's own footprint: ScitiX operates Tier III+ (T3+) data centers in North America, and expects to complete its SOC 2 Type II certification by the end of 2026 - so the security story extends from your API key all the way down to the facilities your models run in. The takeaway. Boson AI's TTS and ASR give you a complete voice loop - expressive, natural speech and accurate, multilingual transcription - behind one OpenAI-compatible API. ScitiX Model Inference makes that loop simple to adopt, fast to scale, and secure by default, from your API key all the way down to the data center. Whether you're giving a product a natural-sounding voice or transcribing speech across 90+ languages, both models are live on ScitiX today - ready the moment you are.