Simplify Logo
Preference Model

Preference Model

Builds RL environments to automate ML

Member of Technical Staff - Research & Post-training

Full-Time
$200k - $350k/yr

+ Equity compensation + 401(k) match

Mid
San Francisco, CA, USA
In Person

On-site in San Francisco.

H1B Sponsorship Available

About the job

Requirements
  • Experience running end-to-end large language model post-training pipelines for models of at least 7B parameters.
  • Proficiency in Python and PyTorch or JAX.
  • Experience with at least one modern reinforcement learning training framework.
  • Experience building and operating machine learning infrastructure at scale.
Responsibilities
  • Train and evaluate models on proprietary reinforcement learning environments to validate data quality, identify gaps in task coverage, and connect environment design with model capability.
  • Architect and optimize reinforcement learning training infrastructure, including training abstractions and distributed experiment management, using frameworks such as Verl, OpenRLHF, or similar.
  • Design, implement, and test training environments, evaluations, and methodologies for reinforcement learning agents.
  • Profile and optimize training runs end-to-end, from data loading through reward computation, to maximize experiment throughput and shorten the research iteration cycle.
Desired Qualifications
  • Experience evaluating model outputs and building reward or evaluation signals.
  • Staying current on post-training research and translating research papers into running code.
  • Strong opinions, loosely held, about structuring reinforcement learning training code for reproducibility and fast iteration.
  • Ability to balance research exploration with engineering rigor.
  • Strong systems design and communication skills.

About the company

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The 2026-08-16 envindex profile says commercial preference and reward data is already offered.
  • South Park Commons backed the 2025 pre-seed, signaling strong early investor validation.
  • If frontier AI training shifts toward RL, demand for specialized environments rises quickly.

What critics are saying

  • Caplight shows only $400,000 raised and a $5.71 million valuation, constraining runway.
  • LinkedIn listed only 11-50 employees on 2026-08-31, limiting product breadth and sales capacity.
  • Frontier labs can build in-house RL environments, compressing Preference Model’s differentiation within 12 months.

What makes Preference Model unique

  • Preference Model’s 2026-07-28 site targets real-world RL environments for ML research automation.
  • Founders claim Anthropic, Stripe, and Datology experience, strengthening technical credibility.
  • The company sells frontier-lab-focused reward and preference infrastructure, not generic model tooling.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%