Full-Time

Member of Technical Staff

Cybersecurity Capabilities

Preference Model

Preference Model

11-50 employees

Builds RL environments to automate ML

Compensation Overview

$180k - $300k/yr

+ Equity

Seattle, WA, USA + 2 more

More locations: Toronto, ON, Canada | San Francisco, CA, USA

In Person

In-person roles; relocation and visa sponsorship available.

Category
IT & Security (1)
Software Engineering (1)
Required Skills
Rust
Python
C/C++

Get referred to Preference Model

Find people who can refer or advise you

Requirements
  • Strong security fundamentals and broad interests across both offensive and defensive work.
  • Hands-on experience finding, exploiting, or patching real vulnerabilities through CTFs, bug bounty work, security research, red/blue team engagements, or shipped security work in industry.
  • Proficiency in Python and systems programming, plus working comfort in at least one low-level language (C, C++, Rust) and one web/application stack.
  • Familiarity with security tooling: fuzzers, sanitizers, debuggers, and disassemblers.
  • Problem solvers who take ownership and drive solutions end-to-end.
  • Passion for staying current with the rapidly evolving security and ML landscape.
  • Ability to meet throughput expectations and respond quickly to feedback.
Responsibilities
  • Design and build RL environments and reward functions that produce clean, learnable signals for frontier models on offensive and defensive security tasks across diverse programming languages.
  • Build environments covering the full vulnerability lifecycle: discovery in source code, exploiting, patching.
  • Build environments for reverse engineering tasks across binaries, bytecode, and obfuscated code.
  • Construct verifiable reward signals using fuzzers, sanitizers, symbolic execution, static analyzers, exploit-success checks, and patch-correctness validation.
  • Collaborate with others to brainstorm and create new ideas and tools to improve the environment building process.
Desired Qualifications
  • Published security research, CVEs, or notable bug bounty findings.
  • Strong CTF background or competitive results at events like DEF CON CTF, or similar.
  • Deep expertise in a specific area: binary exploitation, kernel security, browser/V8 internals, hypervisor security, cryptographic implementation, web application security, or cloud/container security.
  • Experience building or contributing to fuzzing infrastructure, vulnerability scanners, or automated program analysis tools.
  • Experience with ML for code or security.
  • You have built complex interactive RL environments, agent harnesses, or sandboxed evaluation infrastructure.

Preference Model builds RL environments that automate ML research and engineering. What it does: creates reinforcement learning environments that allow researchers and engineers to test and automate tasks in ML workflows. How it works: users interact with programmable environments where an RL agent can perform actions to progress ML experiments, with defined rewards, observations, and interfaces that map to common ML tasks like model training, hyperparameter tuning, or data processing. These environments can be run to automate repetitive research tasks and evaluate ideas at scale. How it differs from competitors: instead of offering general RL tools alone, it targets the automation of ML research and engineering processes, packaging ML tasks into reusable, standardized environments that streamline experimentation and comparison. Goal: to speed up ML research and engineering by providing ready-to-use, reusable RL environments that automate routine experiments and evaluations.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2025

Get referred to Preference Model

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Preference Model solves the bottleneck of lacking high-quality RL training environments for ML research automation.
  • The company partners with frontier AI labs to develop capabilities for next-generation large language models.
  • Preference Model delivers high-impact AI training data solutions operating as a focused San Francisco team.

What critics are saying

  • Anthropic's recruitment of the founding team creates immediate knowledge drain and competitor alignment within three to six months.
  • Frontier AI lab partners building proprietary RL environments negates Preference Model's unique value within six to twelve months.
  • Stripe launching an internal ML research automation tool eliminates cross-industry differentiation within nine to fifteen months.

What makes Preference Model unique

  • Preference Model automates ML research by building diverse RL environments with robust rewards reflecting real-world complexity.
  • The founding team brings prior data and infrastructure experience from Anthropic, Stripe, and Datology AI labs.
  • Preference Model targets the critical need to teach AI models ML research, addressing brittleness in frontier models.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Vision Insurance

Dental Insurance

401(k) Company Match

Company Equity

Meal Benefits

Relocation Assistance