Full-Time

Research Engineer

HUD

HUD

11-50 employees

AI agent evaluation platform and benchmarks

No salary listed

H1B Sponsorship Available

San Francisco, CA, USA

In Person

Hybrid ATS listings conflict with the logistics section, which states the role is on-site in the San Francisco Bay Area. Visa and relocation support are available for strong full-time candidates.

Category
AI & Machine Learning (1)
Required Skills
Python
Data Engineering
Docker
Linux/Unix
Reinforcement Learning

Get referred to HUD

See people who can refer or advise you

Requirements
  • Proficiency in Python, Docker, and Linux environments.
  • Experience working on benchmarks and evaluations, including reasoning about realistic tasks, reliable rubrics, usable environments, and useful trajectories for reinforcement learning training.
  • Strong attention to detail and the ability to spot subtle inconsistencies in data, model behavior, or task design.
  • Experience building tools, pipelines, experiments, or infrastructure without a fully prescribed roadmap.
  • Early-stage startup experience and the ability to work independently in fast-paced environments.
Responsibilities
  • Build systems for creating, running, evaluating, and improving agent training environments.
  • Design experiments to understand model behavior, agent failure modes, and data quality issues.
  • Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops.
  • Work across the full lifecycle of agent training data, including task design, environment setup, trajectory collection, evaluation, and validation.
  • Partner with external vendors to identify bottlenecks and improve the quality and throughput of HUD’s data engine.
  • Build metrics and analyses to determine whether tasks, environments, and evaluations are useful for training frontier agents.
Desired Qualifications
  • Experience building internal tools, research infrastructure, or data pipelines.
  • Experience designing metrics and validation workflows.
  • A background in competitive programming, Olympiad medaling, research, or unusually strong independent project experience.
  • Strong communication skills for remote collaboration across time zones.

HUD provides an evaluation platform for AI agents that perform computer-use tasks. It offers an interface that connects to HUD evaluation environments, allowing users to run benchmarks across hundreds of environments and thousands of tasks. Users can integrate their agents using various adapters, and interact with the system through an asynchronous API for efficient experiments. The platform then collects telemetry and benchmarking results to inform agent improvements. HUD differentiates itself by concentrating on evaluating and improving knowledge-work agents across a wide range of tasks and environments, rather than just building models, and by enabling scalable, tool-enabled testing via adapters and an asynchronous workflow. The company’s goal is to help organizations and developers assess, compare, and enhance their AI agents so they perform better in real-world knowledge-work scenarios.

Company Size

11-50

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2025

Get referred to HUD

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • HUD raised a $16M Series A led by Standard Capital in June 2026.
  • HUD says over 50 businesses use the platform, including vendors selling millions monthly.
  • Native GitHub Agentic Workflows support expands HUD into coding-agent evaluation workflows.

What critics are saying

  • OpenAI and Anthropic can internalize evals, shrinking HUD's core market quickly.
  • HUD faces open-source imitation from Cua, OSWorld, and other benchmark builders.
  • If labs standardize on in-house benchmarks by 2027, HUD becomes a commodity layer.

What makes HUD unique

  • HUD sells live computer-use evals against real software, not static datasets.
  • Its CUA Evals framework targets browser agents, desktop tasks, and RL training pipelines.
  • HUD also runs a vendor marketplace, connecting post-training suppliers directly with AI labs.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Paid Vacation

Paid Holidays

Commuter Benefits

Growth & Insights

Headcount

6 month growth

-8%

1 year growth

-8%

2 year growth

-8%