Full-Time

Machine Learning Engineer New Grad

Machine Learning Capabilities

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$165k - $200k/yr

+ Equity compensation + 401(k) match + Relocation support

H1B Sponsorship Available

Seattle, WA, USA + 1 more

More locations: San Francisco, CA, USA

In Person

Master's, PhD

Category
AI & Machine Learning (1)
Software Engineering (1)
Required Skills
Microsoft Azure
Python
Neural Networks
CUDA
PyTorch
Machine Learning
Infrastructure as Code (IaC)
AWS
NumPy
Reinforcement Learning
Google Cloud Platform

Get referred to Preference Model

See people who can refer or advise you

Requirements
  • Strong machine learning fundamentals and broad research interests, including the ability to understand research deeply and translate it into reinforcement learning from verifiable rewards problems.
  • Expert knowledge in an active deep learning or machine learning research area, with publications or public code demonstrating expertise.
  • Deep understanding of transformer internals.
  • Proficiency in Python, NumPy, and systems programming; PyTorch or JAX proficiency is advantageous.
  • Ability to take ownership and drive solutions end-to-end.
  • Passion for staying current with the rapidly evolving machine learning infrastructure landscape.
  • Ability to meet throughput expectations and respond quickly to feedback.
Responsibilities
  • Design and build reinforcement learning environments and reward schemes that produce clean, learnable signals for frontier models on machine learning research and engineering tasks.
  • Build deep expertise across the frontier of machine learning research, training, and inference infrastructure.
  • Collaborate with others to brainstorm and create new ideas and tools to improve the environment-building process.
  • Stay up to date with the latest research, develop novel approaches, realize them in code, conduct experiments and evaluations, deliver work into production training runs, and collaborate with researchers and engineers.
Desired Qualifications
  • Strong expertise in kernel development, including CUDA, Triton, or Pallas, and optimizing non-trivial neural modules for specific hardware.
  • Research projects, coursework, or personal work involving reinforcement learning environments.
  • Open-source contributions to machine learning infrastructure or reinforcement learning tooling.
  • Experience with a cloud platform such as AWS, GCP, or Azure, or with infrastructure-as-code tools.

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The website was updated May 29, 2026, showing active product development.
  • It posted multiple MTS roles and new-grad hiring updated August 4, 2026.
  • Visa sponsorship and $165k-$350k pay attract scarce ML talent fast.

What critics are saying

  • OpenAI, Anthropic, and Google can internalize these RL environments by 2027.
  • Training-environment tooling commoditizes quickly; customers switch on benchmark wins and price.
  • If frontier labs stop buying external environments, Preference Model loses its entire market.

What makes Preference Model unique

  • Founded 2025, Preference Model builds RL environments for ML research automation.
  • Team includes former Anthropic, Stripe, and Datology operators with infra experience.
  • Company says it partners with frontier AI labs on next-generation LLM capabilities.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance