Full-Time

Lead Research Engineer

Data Quality

Updated on 9/10/2026

HUD

HUD

11-50 employees

AI agent evaluation platform and benchmarks

No salary listed

H1B Sponsorship Available

San Francisco, CA, USA

In Person

On-site only in the San Francisco Bay Area or Singapore offices.

Category
AI & Machine Learning (1)
Required Skills
Python
Data Visualization
Machine Learning
Docker
Quality Assurance (QA)
Linux/Unix

Get referred to HUD

See people who can refer or advise you

Requirements
  • Advanced proficiency in Python, Docker, and Linux environments.
  • Deep intuition for data quality, including the ability to reason about what makes tasks realistic, learnable, diverse, reliable, and useful for training.
  • Experience building quality-control systems, evaluations, benchmarks, synthetic-data pipelines, validation workflows, or model-evaluation infrastructure.
  • Comfort working across messy human and technical systems, including domain experts, vendors, generated data, model outputs, graders, and infrastructure.
  • Strong written communication and the ability to explain methodology clearly to researchers, engineers, labs, and external audiences.
Responsibilities
  • Lead HUD’s data quality strategy, including building quality-control systems, defining and enforcing quality standards, and designing experiments to grade agent outputs.
  • Develop methods for validating synthetic data at scale, including failure-mode analysis, task-mutation checks, and trajectory auditing.
  • Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data-generation workflows.
  • Turn qualitative research insights into production systems, internal tools, dashboards, validation pipelines, and feedback loops.
  • Build internal research judgment around what makes agent-training data useful rather than superficially correct.
  • Mentor other research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
Desired Qualifications
  • Experience leading teams on ambiguous technical projects from problem definition through implementation and iteration.
  • Experience working with subject-matter experts to capture domain judgment and convert it into scalable review or generation systems.
  • Comfort designing metrics, experiments, and quality-assurance and quality-control processes rather than only executing them.
  • Early-stage startup experience and the ability to work independently in fast-paced environments.
  • Attention to detail and the ability to identify subtle inconsistencies or edge cases in data.

HUD provides an evaluation platform for AI agents that perform computer-use tasks. It offers an interface that connects to HUD evaluation environments, allowing users to run benchmarks across hundreds of environments and thousands of tasks. Users can integrate their agents using various adapters, and interact with the system through an asynchronous API for efficient experiments. The platform then collects telemetry and benchmarking results to inform agent improvements. HUD differentiates itself by concentrating on evaluating and improving knowledge-work agents across a wide range of tasks and environments, rather than just building models, and by enabling scalable, tool-enabled testing via adapters and an asynchronous workflow. The company’s goal is to help organizations and developers assess, compare, and enhance their AI agents so they perform better in real-world knowledge-work scenarios.

Company Size

11-50

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2025

Get referred to HUD

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • HUD raised a Series A on June 18, 2026, signaling investor conviction.
  • Over 50 businesses already use HUD, proving real commercial demand.
  • HUD’s platform exports graded traces and supports training data, expanding monetization beyond benchmarks.

What critics are saying

  • Scale and Mercor now attack computer-use environments; HUD faces faster-funded rivals in 2026.
  • HUD’s fast-moving v6 API and docs create execution risk for customers and contributors.
  • If frontier labs internalize environment generation, HUD becomes a replaceable tooling layer by 2027.

What makes HUD unique

  • HUD’s v6 platform unifies environments, tasksets, traces, and replayable graded attempts.
  • HUD’s open-source Python SDK and REST API lower integration friction for agents.
  • HUD focuses on computer-use and long-horizon evaluations, a narrower wedge than generic observability.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Paid Vacation

Paid Holidays

Commuter Benefits

Growth & Insights

Headcount

6 month growth

-14%

1 year growth

-14%

2 year growth

-14%