Full-Time

ML Infrastructure Engineer

Post-training

Updated on 8/21/2026

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$180k - $300k/yr

+ Equity compensation + Relocation support

H1B Sponsorship Available

Seattle, WA, USA + 1 more

More locations: San Francisco, CA, USA

In Person

On-site work is required.

Category
DevOps & Infrastructure (1)
Software Engineering (1)
Required Skills
LLM
Kubernetes
Distributed Systems
PyTorch
Machine Learning
Data Engineering
AWS
Reinforcement Learning
Google Cloud Platform

Get referred to Preference Model

See people who can refer or advise you

Requirements
  • Strong software engineering fundamentals and hands-on experience building production-grade large language model inference and training infrastructure, ideally from the ground up.
  • Experience with large language model training and inference internals such as transformers, distributed training, and inference libraries including vLLM, SGLang, and Megatron.
  • Experience working on reinforcement learning training frameworks such as Slime, veRL, Ray Train, and SkyRL.
  • Significant experience and understanding of distributed systems principles.
  • Hands-on experience with cloud platforms such as Amazon Web Services and Google Cloud Platform, and container orchestration with Kubernetes, building systems for high-throughput, low-latency workloads.
  • Experience with data engineering tools and building robust, scalable data pipelines.
  • Proficiency in core machine learning frameworks such as PyTorch or JAX.
  • Ability to balance production rigor with the pace of fast-moving research and communicate infrastructure tradeoffs clearly to researchers who are not infrastructure specialists.
Responsibilities
  • Design, build, and scale the compute, scheduling, and data infrastructure that powers post-training research on in-house reinforcement learning environments.
  • Develop and maintain core machine learning framework primitives and internal tooling that researchers rely on daily, accelerating reproducible experimentation and reducing time from idea to result.
  • Build evaluation and benchmarking infrastructure, monitoring, logging, debugging tooling, and automated testing and deployment systems so failures are caught early and infrastructure remains reliable as it scales.
  • Partner directly with Research Engineers to translate research needs into infrastructure requirements and ship quickly in response to their feedback.

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The website was updated May 29, 2026, showing active product development.
  • It posted multiple MTS roles and new-grad hiring updated August 4, 2026.
  • Visa sponsorship and $165k-$350k pay attract scarce ML talent fast.

What critics are saying

  • OpenAI, Anthropic, and Google can internalize these RL environments by 2027.
  • Training-environment tooling commoditizes quickly; customers switch on benchmark wins and price.
  • If frontier labs stop buying external environments, Preference Model loses its entire market.

What makes Preference Model unique

  • Founded 2025, Preference Model builds RL environments for ML research automation.
  • Team includes former Anthropic, Stripe, and Datology operators with infra experience.
  • Company says it partners with frontier AI labs on next-generation LLM capabilities.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance