Simplify Logo
Preference Model

Preference Model

Builds RL environments to automate ML

Machine Learning Infrastructure Engineer - Post-training

Full-Time
$200k - $350k/yr

+ Equity compensation + Relocation support

Senior
San Francisco, CA, USA
In Person

On-site in San Francisco.

H1B Sponsorship Available

About the job

Requirements
  • Strong software engineering fundamentals and hands-on experience building production-grade large language model inference and training infrastructure, ideally from the ground up.
  • Experience building large language model training and inference internals such as transformers and distributed training, and working with inference libraries such as vLLM, SGLang, or Megatron.
  • Experience working with reinforcement learning training frameworks such as Slime, veRL, Ray Train, or SkyRL.
  • Significant experience and understanding of distributed systems principles, with hands-on experience using cloud platforms such as AWS or Google Cloud Platform and container orchestration with Kubernetes to build high-throughput, low-latency workloads.
  • Experience with data engineering tools and building robust, scalable data pipelines.
  • Proficiency in core machine learning frameworks such as PyTorch or JAX.
  • Ability to balance production rigor with the pace of fast-moving research and communicate infrastructure tradeoffs clearly to researchers who are not infrastructure specialists.
Responsibilities
  • Design, build, and scale the compute, scheduling, and data infrastructure that powers post-training research on in-house reinforcement learning environments.
  • Develop and maintain core machine learning framework primitives and internal tooling that researchers rely on daily, accelerating reproducible experimentation and reducing time from idea to result.
  • Build evaluation and benchmarking infrastructure, monitoring, logging, debugging tooling, and automated testing and deployment systems so failures are caught early and infrastructure remains reliable as it scales.
  • Partner directly with Research Engineers to translate research needs into infrastructure requirements and ship quickly in response to their feedback.

About the company

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The 2026-08-16 envindex profile says commercial preference and reward data is already offered.
  • South Park Commons backed the 2025 pre-seed, signaling strong early investor validation.
  • If frontier AI training shifts toward RL, demand for specialized environments rises quickly.

What critics are saying

  • Caplight shows only $400,000 raised and a $5.71 million valuation, constraining runway.
  • LinkedIn listed only 11-50 employees on 2026-08-31, limiting product breadth and sales capacity.
  • Frontier labs can build in-house RL environments, compressing Preference Model’s differentiation within 12 months.

What makes Preference Model unique

  • Preference Model’s 2026-07-28 site targets real-world RL environments for ML research automation.
  • Founders claim Anthropic, Stripe, and Datology experience, strengthening technical credibility.
  • The company sells frontier-lab-focused reward and preference infrastructure, not generic model tooling.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%