Preference Model builds reinforcement learning environments and harnesses for training frontier AI models, focusing on how reward signals shape model behavior. It provides simulation environments and reward/graders that guide learning, testing and hardening signals across millions of rollouts while including fail-safes to prevent unsafe behavior. The company emphasizes targeting reward design and alignment bottlenecks, with a team from Anthropic, Stripe, Google DeepMind and others, backed by notable investors. Its goal is to diffuse AI capabilities widely without concentrating power, while keeping strong safeguards to keep frontier systems safe.
Company Size
11-50
Company Stage
Seed
Total Funding
$16M
Headquarters
N/A
Founded
2025
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Vision Insurance
Dental Insurance
401(k) Company Match
Company Equity
Meal Benefits
Relocation Assistance
On October 7, 2026, Preference Model, a superintelligence data research company, emerged from stealth and announced the completion of a $16 million seed financing led by a16z. SignalFire, South Park…