Simplify Logo
Vmax

Vmax

Automates generation of reinforcement learning environments

Member of Technical Staff - RL Infrastructure

Full-Time
$300k - $500k/yr
Senior
San Francisco, CA, USA
Hybrid

Hybrid role; requires some on-site presence in San Francisco, California.

About the job

Requirements
  • Strong software engineering experience.
  • Experience building infrastructure for LLM inference and/or RL training.
  • Experience with GPU clusters, distributed training, model serving, or high-throughput inference systems.
  • Familiarity with vLLM, SGLang and modern LLM-RL training frameworks
  • Strong understanding of system reliability, observability, testing, debugging, and performance optimization.
  • Ability to work closely with ML researchers and translate messy experimental workflows into durable infrastructure.
  • Experience building tools, platforms, or services used by other technical users.
  • Strong judgment around technical tradeoffs: when to prototype, when to harden, when to simplify, and when to redesign.
  • Clear written and verbal communication, especially around system design, operational risks, and engineering tradeoffs.
Responsibilities
  • Build infrastructure for distributed RL training and inference across thousands of GPUs
  • Improve the reliability, debuggability, and throughput of RL experiments.
  • Build interfaces that allow researchers and applied ML engineers to launch, inspect, compare, and reproduce experiments easily.
  • Own infrastructure projects end to end, from architecture and implementation through deployment, documentation, and long-term maintenance.
  • Identify and eliminate bottlenecks in training, rollout generation, eval execution, data movement, and cluster utilization.
  • Maintain engineering standards for RL infrastructure, including testing, observability, versioning, and reproducibility.
Desired Qualifications
  • Experience supporting research teams or fast-moving ML teams.
  • Experience at a high engineering bar organization where reliability, ownership, and code quality were central.
  • Evidence of strong independent technical work, such as open-source projects, infrastructure projects, competitions, or substantial systems built from scratch.
  • Experience reducing operational complexity in systems that had become brittle, slow, or hard to debug.

About the company

Vmax.ai builds tools to automate reinforcement learning (RL) development. It creates scalable RL environments from proprietary data so engineers can train agents for long-horizon tasks without a lot of manual RL engineering. The core product concept is to automatically transform company data and evaluation metrics into reusable RL environments, enabling post-training of large language model–based agents for domain-specific use cases. The company distinguishes itself by combining automated environment design with data-driven RL environment generation, aiming to cut human intervention in RL workflows. The founding team’s deep RL background, stealth mode status, and backing from South Park Commons position it to pursue enterprise-grade RL automation and domain-specific AI through a platform that handles data-to-environment conversion and subsequent agent fine-tuning. Its goal is to scale RL development by reducing setup work and enabling long-horizon tasks across domains.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2025

Get referred to Vmax

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • South Park Commons still lists Vmax active and hiring in August 2026.
  • Vmax published unix-ctf and PROPEL in 2026, proving real research velocity.
  • Hiring for RL Infrastructure and Applied RL suggests customer-ready systems work, not just theory.

What critics are saying

  • Open-ended learning remains unproven commercially; research outputs do not equal a durable product.
  • $300k-$500k roles in 2026 signal heavy burn before revenue arrives.
  • OpenAI, Anthropic, and Nvidia can commoditize RL environment tooling, killing Vmax's moat.

What makes Vmax unique

  • Vmax builds proprietary-data-to-RL-environment pipelines, not generic model fine-tuning.
  • Its unix-ctf and PROPEL research show environment generation for long-horizon agents.
  • Matthew Sargent and Augustine Mavor-Parker bring PhD-level RL depth from UCL, Redwood, and Illumina.

Help us improve and share your feedback! Did you find this helpful?

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%