Full-Time

Machine Learning Engineer

Machine Learning Capabilities

Posted on 9/12/2026

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$200k - $350k/yr

+ Equity compensation + Relocation support

H1B Sponsorship Available

San Francisco, CA, USA

In Person

On-site role in San Francisco.

Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
Neural Networks
CUDA
PyTorch
Machine Learning
Reinforcement Learning

Get referred to Preference Model

See people who can refer or advise you

Requirements
  • At least 5 years of experience working in machine learning or research, primarily on large language models and transformer models.
  • Strong machine learning fundamentals and broad research interests, with the ability to understand technical topics deeply and translate them into reinforcement learning with verifiable rewards problems.
  • Proficiency in Python and systems programming, plus proficiency in at least one of PyTorch or JAX.
  • Ability to take ownership and drive solutions end-to-end.
  • Passion for staying current with the rapidly evolving machine learning infrastructure landscape.
  • Ability to meet throughput expectations and respond quickly to feedback.
Responsibilities
  • Design and build reinforcement learning environments and reward functions that produce clean, learnable signals for frontier models on machine learning research and engineering tasks.
  • Build deep expertise across the frontier of machine learning research, training, and inference infrastructure.
  • Collaborate with others to brainstorm and create new ideas and tools to improve the environment-building process.
Desired Qualifications
  • Expert knowledge in an active deep learning or machine learning research area, with publications or public code demonstrating that expertise.
  • Research experience and a PhD or Master of Science degree.
  • Deep understanding of transformer internals and training and inference of modern large language models, with experience using inference libraries such as vLLM or SGLang.
  • Strong expertise in kernel development using CUDA, Triton, or Pallas.
  • Experience building complex interactive reinforcement learning environments.

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The 2026-08-16 envindex profile says commercial preference and reward data is already offered.
  • South Park Commons backed the 2025 pre-seed, signaling strong early investor validation.
  • If frontier AI training shifts toward RL, demand for specialized environments rises quickly.

What critics are saying

  • Caplight shows only $400,000 raised and a $5.71 million valuation, constraining runway.
  • LinkedIn listed only 11-50 employees on 2026-08-31, limiting product breadth and sales capacity.
  • Frontier labs can build in-house RL environments, compressing Preference Model’s differentiation within 12 months.

What makes Preference Model unique

  • Preference Model’s 2026-07-28 site targets real-world RL environments for ML research automation.
  • Founders claim Anthropic, Stripe, and Datology experience, strengthening technical credibility.
  • The company sells frontier-lab-focused reward and preference infrastructure, not generic model tooling.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%