Full-Time

Member of Technical Staff

Head of Quality

Updated on 9/11/2026

Plato

Plato

11-50 employees

Delivers reproducible AI training environments

Compensation Overview

$180k - $280k/yr

San Francisco, CA, USA

In Person

Category
QA & Testing (1)
Required Skills
Python
Machine Learning
Quality Assurance (QA)
Reinforcement Learning

Get referred to Plato

See people who can refer or advise you

Requirements
  • At least 3 years of experience in software engineering, machine learning engineering, or research systems, with proficiency in modern programming languages such as Python.
  • Understanding of how optimizers exploit edge cases, including how agents can game reward functions, break sandboxes, or fake completion.
  • Ability to analyze raw rollouts, agent reasoning traces, and verification code to identify subtle ungrounded assumptions or hallucinated logic.
  • Experience building or leading a technical evaluation, quality assurance, or data verification function from zero to one.
  • Ability to maintain delivery quality standards under intense customer pressure and recognize that contaminated data is more harmful than delayed delivery.
Responsibilities
  • Own final sign-off before environments and datasets are delivered to frontier laboratories, auditing trajectories, tasks, and reward dynamics for correctness, feasibility, and signal density.
  • Build automated red-teaming suites to stress-test task feasibility and verifier integrity, eliminating reward hacking, grader tampering, and impossible task traps.
  • Architect judge models, sandbox replay harnesses, and rollout forensics to detect synthetic drift, out-of-distribution behaviors, and leaky states upstream.
  • Hire and direct a team of quality assurance engineers, domain specialists, and technical reviewers, combining automated agentic checks with human-in-the-loop review.
  • Translate downstream model failure modes into concrete generator constraints so defects are prevented at the generation stage rather than caught during review.
Desired Qualifications
  • Experience with reinforcement learning training dynamics, automated large language model evaluations, agent sandboxing, or synthetic trajectory generation.
  • Background designing adversarial test suites, code execution verifiers, or formal verification systems.
  • Experience handling client-facing technical evaluations and failure postmortems with frontier artificial intelligence research teams.

Plato creates datasets and simulated environments to train and evaluate AI web agents. It produces static, reproducible environments from human demonstrations, providing labs and agent developers a controlled way to benchmark and reinforce-behavior learning on real-world web tasks. Unlike other providers, Plato focuses on supplying the data and virtual “gyms” that power modern AI progress in reinforcement learning. Its goal is to help advance AI by making reliable, testable environments and benchmarks for web-based tasks.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2024

Get referred to Plato

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Plato raised $14.5 million in February 2026, signaling investor confidence and runway.
  • RL Engineering listed Plato as an active confirmed vendor in September 2026.
  • The browser-agent market needs evaluation environments for RL training and benchmarking right now.

What critics are saying

  • OpenAI and Anthropic keep bundling native computer-use capabilities, compressing Plato's pricing by 2027.
  • Plato's replicas need constant upkeep as websites change, inflating engineering costs and breaking benchmarks.
  • A larger labs-owned environment stack would make Plato a replaceable vendor, erasing standalone demand.

What makes Plato unique

  • Plato sells reproducible web-agent gyms, not just scraped datasets, in 2026.
  • Its Python SDK and dedicated tenants support controlled browser and Linux-desktop evaluation.
  • It recreates Amazon, Airbnb, Gmail-style tasks with scoring and state tracking.

Help us improve and share your feedback! Did you find this helpful?

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%