Full-Time

Agent Harness Engineer

Model AI

Model AI

No salary listed

Palo Alto, CA, USA

In Person

Category
AI & Machine Learning (1)
Required Skills
Python
Observability
DevOps
Reinforcement Learning
Requirements
  • Strong Python engineering skills.
  • Experience building large language model agents, evaluation harnesses, developer tools, applied artificial intelligence systems, or automation workflows.
  • Strong systems instincts and the ability to build reliable, debuggable tooling.
  • Experience with testing, benchmarking, continuous integration, code analysis, or automated software workflows.
  • Ability to reason about agent behavior, failure modes, evaluation quality, and product usefulness.
  • Comfort working across agent logic, backend systems, infrastructure, and user-facing product requirements.
  • Strong product intuition and the ability to build demos that can evolve into real products.
  • High ownership and the ability to operate effectively in an early-stage startup environment.
  • Hands-on technical excellence and strong engineering judgment.
  • End-to-end ownership, from design to implementation to production outcomes.
  • Ability to do deep, focused work and sustain execution.
  • Clear communication with teammates, customers, and stakeholders.
  • Comfort with ambiguity, rapid change, and wearing multiple hats.
  • Low ego, high integrity, high accountability, and strong collaboration.
  • Continuous learning and a belief that judgment, intelligence, and capability compound over time.
Responsibilities
  • Build agent harnesses for multi-agent workflows and ralph loops.
  • Design evaluation environments to measure agent performance, reliability, cost, latency, and quality.
  • Build testing systems that connect agent behavior, code changes, and evaluation results.
  • Create workflows where agents can inspect code, make changes, run tests, and determine whether performance improved.
  • Develop shared-memory, coordination, and task-management systems for teams of agents.
  • Build infrastructure for tracking task success, regressions, cost, latency, quality, and long-term progress.
  • Work on applied coding-agent workflows, including refactoring, testing, debugging, code review, and technical-debt reduction.
  • Support customer demos and applied workflows that turn agent research into usable product experiences.
  • Collaborate closely with machine learning systems engineers to make agent workloads run efficiently on Agent Cloud.
  • Help turn research prototypes into reliable, measurable, production-quality systems.
Desired Qualifications
  • Experience with coding agents, evaluation harnesses, SWE-bench-style environments, or tool-use systems.
  • Experience with multi-agent or long-running agent workflows.
  • Experience with large language model observability, tracing, evaluations, or reinforcement learning workflows.
  • Experience with continuous integration/continuous delivery, automated testing, static analysis, or developer tooling.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A