P

Preference Model

Reinforcement learning environments for frontier AI

Member of Technical Staff - Research Engineer, Post-training

Full-TimeUpdated on 10/6/2026
$200k - $350k/yr+ Equity compensation + 401(k) match
Mid
San Francisco, CA, USA
In Person
H1B Sponsorship Available

About the job

Requirements
  • Experience running end-to-end post-training pipelines for large language models at least 7 billion parameters in size.
  • Proficiency in Python and PyTorch or JAX.
  • Experience with at least one modern reinforcement learning training framework.
  • Experience building and operating machine learning infrastructure at scale.
Responsibilities
  • Train and evaluate models on proprietary reinforcement learning environments to validate data quality, surface task-coverage gaps, and connect environment design to model capability.
  • Architect and optimize reinforcement learning training infrastructure, including training abstractions and distributed experiment management, using frameworks such as Verl, OpenRLHF, or similar; scale systems for increasingly complex research workflows.
  • Design, implement, and test training environments, evaluations, and methodologies for reinforcement learning agents.
  • Profile and optimize training runs end to end, from data loading through reward computation, to maximize experiment throughput and shorten research iteration cycles.
Desired Qualifications
  • Experience evaluating model outputs and building reward or evaluation signals.
  • Stay current on post-training research and translate papers into running code.
  • Strong opinions, held loosely, about structuring reinforcement learning training code for reproducibility and fast iteration.
  • Ability to balance research exploration with engineering rigor.
  • Strong systems design and communication skills.

About the company

Preference Model builds reinforcement learning environments and harnesses for training frontier AI models, focusing on how reward signals shape model behavior. It provides simulation environments and reward/graders that guide learning, testing and hardening signals across millions of rollouts while including fail-safes to prevent unsafe behavior. The company emphasizes targeting reward design and alignment bottlenecks, with a team from Anthropic, Stripe, Google DeepMind and others, backed by notable investors. Its goal is to diffuse AI capabilities widely without concentrating power, while keeping strong safeguards to keep frontier systems safe.

Company Size

11-50

Company Stage

Seed

Total Funding

$16M

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The website updated October 3, 2026, shows active product focus and current momentum.
  • Preference Model is hiring across five engineering roles with $200K–$350K salaries.
  • South Park Commons listed Jennifer Zhou’s startup, signaling credible early network validation.

What critics are saying

  • Preference Model has only 17 employees and five open San Francisco roles on September 25, 2026.
  • Caplight shows just $400,000 pre-seed raised on July 1, 2025, limiting runway.
  • OpenAI, Anthropic, and Google can internalize RL environments, erasing Preference Model’s standalone market.

What makes Preference Model unique

  • Preference Model builds RL environments for ML research, not generic training datasets.
  • The team includes ex-Anthropic and ex-Stripe operators, plus South Park Commons backing.
  • Their low-level kernel environments target CUDA, accelerators, and hardware-specific benchmarks.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance

Growth & Insights and Company News

Headcount

6 month growth

↑ 14%

1 year growth

↑ 14%

2 year growth

↑ 14%
Wilson Sonsini Goodrich & Rosati
Oct 7th, 2026
Wilson Sonsini Advises Preference Model on $16 Million Seed Funding as it Emerges from Stealth

On October 7, 2026, Preference Model, a superintelligence data research company, emerged from stealth and announced the completion of a $16 million seed financing led by a16z. SignalFire, South Park…