Full-Time

Member of Technical Staff New Grad

Machine Learning Capabilities

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$165k - $200k/yr

+ Equity

H1B Sponsorship Available

Seattle, WA, USA + 2 more

More locations: Toronto, ON, Canada | San Francisco, CA, USA

In Person

Relocation and visa sponsorship available.

Category
AI & Machine Learning (1)
Software Engineering (1)
Required Skills
Python
Pytorch
Machine Learning
Reinforcement Learning

Get referred to Preference Model

Find people who can refer or advise you

Requirements
  • You have strong ML fundamentals and broad research interests.
  • You have proficiency in Python and systems programming; ideally PyTorch or JAX.
  • You are a smart problem solver who takes ownership and drives solutions end-to-end.
  • You have a passion for staying current with the rapidly evolving ML infrastructure landscape.
  • You have the ability to meet throughput expectations and respond quickly to feedback.
Responsibilities
  • Design and build reinforcement learning environments and reward schemes that produce clean, learnable signals for frontier models on machine learning research and engineering tasks.
  • Build deep expertise across the frontier of machine learning research, training, and inference infrastructure.
  • Collaborate with others to brainstorm and create new ideas and tools to improve the environment building process.
Desired Qualifications
  • Expert knowledge in an active deep learning/machine learning research area, with publications or public code to show for it. Research experience (PhD, MS) is a big plus.
  • Deep understanding of transformer internals.
  • Strong expertise in kernel development (CUDA, Triton, Pallas), optimizing non-trivial neural modules to specific hardware.
  • Research projects, coursework, or personal work involving RL environments (any framework, any scale).
  • Open-source contributions to ML infrastructure or RL tooling.
  • Experience with any cloud platform (Amazon Web Services, Google Cloud Platform, Microsoft Azure) or infrastructure-as-code tools

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Preference Model solves the bottleneck of lacking high-quality RL training environments for ML research automation.
  • The company partners with frontier AI labs to develop capabilities for next-generation large language models.
  • Preference Model delivers high-impact AI training data solutions operating as a focused San Francisco team.

What critics are saying

  • Anthropic's recruitment of the founding team creates immediate knowledge drain and competitor alignment within three to six months.
  • Frontier AI lab partners building proprietary RL environments negates Preference Model's unique value within six to twelve months.
  • Stripe launching an internal ML research automation tool eliminates cross-industry differentiation within nine to fifteen months.

What makes Preference Model unique

  • Preference Model automates ML research by building diverse RL environments with robust rewards reflecting real-world complexity.
  • The founding team brings prior data and infrastructure experience from Anthropic, Stripe, and Datology AI labs.
  • Preference Model targets the critical need to teach AI models ML research, addressing brittleness in frontier models.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance