Full-Time

Member of Technical Staff

Machine Learning Capabilities

Updated on 9/11/2026

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$165k - $200k/yr

+ Equity compensation + 401(k) match

H1B Sponsorship Available

San Francisco, CA, USA

In Person

On-site in San Francisco.

Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
Microsoft Azure
Python
CUDA
PyTorch
Machine Learning
Infrastructure as Code (IaC)
AWS
NumPy
Reinforcement Learning
Google Cloud Platform

Get referred to Preference Model

See people who can refer or advise you

Requirements
  • Strong machine learning fundamentals and broad research interests, including the ability to understand machine learning topics deeply and translate them into reinforcement learning with verifiable rewards problems.
  • Expert knowledge in an active deep learning or machine learning research area, with publications or public code demonstrating expertise.
  • Deep understanding of transformer internals.
  • Proficiency in Python, NumPy, and systems programming.
  • Ability to solve problems, take ownership, and drive solutions end-to-end.
  • Passion for staying current with the rapidly evolving machine learning infrastructure landscape.
  • Ability to meet throughput expectations and respond quickly to feedback.
Responsibilities
  • Design and build reinforcement learning environments and reward schemes that produce clean, learnable signals for frontier models on machine learning research and engineering tasks.
  • Build deep expertise across the frontier of machine learning research, training, and inference infrastructure.
  • Collaborate with others to brainstorm and create new ideas and tools to improve the environment-building process.
  • Develop novel approaches, implement them in code, conduct experiments and evaluations, and deliver work into production training runs.
Desired Qualifications
  • Proficiency with PyTorch or JAX.
  • Strong expertise in kernel development using CUDA, Triton, or Pallas, including optimizing non-trivial neural modules for specific hardware.
  • Research projects, coursework, or personal work involving reinforcement learning environments.
  • Open-source contributions to machine learning infrastructure or reinforcement learning tooling.
  • Experience with a cloud platform such as AWS, Google Cloud Platform, or Azure, or with infrastructure-as-code tools.

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The 2026-08-16 envindex profile says commercial preference and reward data is already offered.
  • South Park Commons backed the 2025 pre-seed, signaling strong early investor validation.
  • If frontier AI training shifts toward RL, demand for specialized environments rises quickly.

What critics are saying

  • Caplight shows only $400,000 raised and a $5.71 million valuation, constraining runway.
  • LinkedIn listed only 11-50 employees on 2026-08-31, limiting product breadth and sales capacity.
  • Frontier labs can build in-house RL environments, compressing Preference Model’s differentiation within 12 months.

What makes Preference Model unique

  • Preference Model’s 2026-07-28 site targets real-world RL environments for ML research automation.
  • Founders claim Anthropic, Stripe, and Datology experience, strengthening technical credibility.
  • The company sells frontier-lab-focused reward and preference infrastructure, not generic model tooling.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%