Full-Time

Machine Learning Engineer

Voice Conversion

Cantina

Cantina

51-200 employees

Subscription-based platform for AI-driven social gaming

Compensation Overview

$200k - $220k/yr

+ Equity

Remote in USA + 1 more

More locations: Europe

Remote

Category
AI & Machine Learning (1)
Required Skills
CUDA
PyTorch
Machine Learning
C/C++

Get referred to Cantina

See people who can refer or advise you

Requirements
  • Exceptional research and development experience with large-scale audio models exceeding 8 billion parameters and 500,000 hours of data.
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning mechanisms, and distillation.
  • Deep hands-on experience training audio variational autoencoders, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
  • Strong experience with multi-node, multi-GPU distributed training using Fully Sharded Data Parallel, DeepSpeed, or an equivalent technology.
  • Strong software engineering skills with a proven track record of building complex systems.
  • Strong experience with PyTorch and performance work, including profiling, CUDA, Triton, and C++ as needed, plus writing reliable production-quality code.
  • Experience shipping large-scale speech, audio, or multimodal generative models to production.
  • Background working with large-scale machine-learning data and iterating on data while triangulating quality using subjective and objective signals.
  • Experience with voice cloning, speech control and steerability, or expressive speech generation.
  • Notable publications and/or open-source contributions in speech, audio, or machine learning.
Responsibilities
  • Architect, implement, pre-train, fine-tune, and post-train or align large-scale speech models, including methods such as GRPO and DPO.
  • Design, run, and analyze scientific experiments to advance understanding of the models.
  • Develop and improve developer tooling to enhance team productivity.
  • Contribute to the entire stack, from low-level optimizations to high-level model design.
  • Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic-data strategies.
  • Design automated objective and subjective evaluations, including listening tests, SV, WER, and ASR-based metrics, robustness and bias checks, and red-team studies.
  • Harden the training-to-evaluation-to-inference pipeline; profile latency, memory, and cost; and meet production service-level agreements with robust monitoring and rollback.
  • Contribute to safety and consent guardrails and misuse and abuse mitigation for responsible speech technology.
Desired Qualifications
  • The posting emphasizes collaboration with research, data, infrastructure, and product teams to ship measurable improvements.
  • The role involves designing listening tests and metrics that correlate with user-perceived quality.

Cantina is an invitation-only platform that blends social gaming with artificial intelligence in a 24/7 virtual club called The Cantina. Members subscribe to access features that let them add AI bots with distinct personalities, chat, play games, and even generate AI art, with bot behavior adapting to user inputs. The service differentiates itself through a private, personalized, dynamic experience that combines human–AI interaction with user-generated content in a shared space. Its goal is to build a steady, subscription-based platform that appeals to tech-savvy users and brands by delivering ongoing, customized digital interactions.

Company Size

51-200

Company Stage

Series B

Total Funding

$33.3M

Headquarters

New York City, New York

Founded

2014

Get referred to Cantina

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Cantina raised $8 million on July 30, 2026, reaching $16.5 million total.
  • August 26, 2026 releases added Claude Code telemetry, MCP inventory, and Slack approvals.
  • Cantina now serves healthcare, fintech, and Fortune 500 customers, expanding enterprise credibility.

What critics are saying

  • Framework Ventures is the only named July 2026 investor, limiting investor diversity.
  • Wiz, SentinelOne, CrowdStrike, and Aikido crowd the same agentic security category.
  • If Clarion misses remediation speed, customers will keep buying point tools instead.

What makes Cantina unique

  • Clarion moves security from discovery to verified resolution, launched July 30, 2026.
  • Apex performs autonomous offensive security, writing fixes and opening pull requests.
  • Anthropic’s Cyber Verification Program access gives Cantina privileged model capability for cyber workflows.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Paid Vacation

Paid Sick Leave

401(k) Retirement Plan

401(k) Company Match

Parental Leave

Fertility Treatment Support

Company Equity

Home Office Stipend

Flexible Work Hours

Hybrid Work Options

Company News

Harro
May 24th, 2026
MC&V adds Canva and Canteen to client roster.

MC&V adds Canva and Canteen to client roster. AI-native production company MC&V has secured two new client wins: Canva and Canteen, adding them to its roster. The wins reflect growing demand from global brands and culturally led organisations for AI production partners with experience delivering quality and efficiency in commercial production environments. The appointments build on existing relationships, with MC&V founders Marie-Céline Merret Wirström and Vinne Schifferstein Vidal having previously worked with Canteen and Canva. AI moves from experimentation to production. The client wins come as brands move AI from experimentation into real campaign production. MC&V said the challenge is no longer simply generating content, but designing production workflows that can hold up commercially, legally and operationally. That includes questions around IP, ownership, approvals, governance and production feasibility. Vidal said the real value now lies in knowing how to turn AI tools into work that can perform for brands and agencies. "Right now, almost everyone has access to the same AI tools, but what's becoming more valuable is knowing how to turn those tools into work that holds up creatively and commercially," Vidal said. "The tools have progressed incredibly quickly, but direction, craft and experience with brand production still matter enormously. That's the thinking behind MC&V. "We're building around people who understand storytelling, production and how to deliver work properly for brands and agencies." New AI director roster. Founded earlier this year, MC&V positions itself as an AI-native production company built around workflow design, creative direction and production systems. The company brings more than two years of experience delivering AI campaigns for brands including Adidas, Uniqlo, KFC and Afterpay. As part of its expansion, MC&V has added AI directors Jodie Heenan, Josef "Seppi" Scholler and Jagger Waters to its creative roster. Heenan is an Australian-based AI artist with a background in design, motion and VFX. She is also Creative Director of her own company and has worked with brands including Cadbury and Twinings. Scholler is known for cinematic visual style, hybrid production and AI-driven world-building, combining emotional scenery with generative workflows. Waters is a multidisciplinary AI creator, writer and producer with more than a decade of experience across film, audio, live events and digital storytelling. Production thinking becomes critical. MC&V co-founder Marie-Céline Merret Wirström said the shift to AI-native workflows is changing what production requires. "Production is becoming less about departments and more about how you design the workflow around the idea and bring the best talent around the brief," Merret Wirström said. "Clients are asking different questions now. Not just what AI can do, but how it fits into their process, their approvals and how it scales. That's where production thinking becomes critical." MC&V will continue to build its offering across hybrid production, AI-assisted creative development, and fully AI-native campaign execution, working with agencies and brands to design production pipelines that balance creative ambition with commercial realities. Top image: (L to R) - Marie-Céline Merret, Jagger Waters, Josef Scholler, Jodie Herman, Vinne Schifferstein.