Full-Time

Senior Backend Engineer

Judgment Labs

Judgment Labs

11-50 employees

Agent behavior monitoring platform for AI

No salary listed

San Francisco, CA, USA

In Person

Category
Software Engineering (1)
Required Skills
RabbitMQ
Airflow
Distributed Systems
Apache Spark
Machine Learning
Apache Kafka
OpenTelemetry
ClickHouse
Next.js
Observability
Data Modeling

Get referred to Judgment Labs

See people who can refer or advise you

Requirements
  • Strong backend engineering experience building and operating production systems under real load.
  • Excellent fundamentals in distributed systems, API design, data modeling, reliability, and performance.
  • Experience working with high-volume event, trace, log, metric, or telemetry data.
  • Strong intuition for data systems, including query patterns, storage layout, indexing, partitioning, latency, correctness, and cost.
  • Ability to debug production issues, improve observability, scale bottlenecks, and clean up abstractions as the product evolves.
  • Ability to work across backend, data, product, and infrastructure boundaries.
  • Product judgment and willingness to ship across the stack when needed.
  • Ability to write design documents, review code changes, explain tradeoffs, and unblock others.
Responsibilities
  • Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics.
  • Own the API surface used by the Judgment platform user interface, software development kits, JudgmentHub libraries, Model Context Protocol server, Slack agent, and customer integrations.
  • Build and operate the RabbitMQ and Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, and tenant-level scheduling.
  • Optimize the ClickHouse online analytical processing layer through schema design, partitioning, skip indexes, full-text-search pruning, query rewrites, deduplication, pagination correctness, and storage growth management.
  • Turn raw spans, conversations, tool calls, scorer outputs, and agent-judge results into clean data models for evaluations, labeling, context engineering, and reinforcement-learning workflows.
  • Ship features end to end across Next.js, backend application programming interfaces, queues and workflows, and the data layer.
  • Work directly with customers to understand agent failures, identify useful data, and structure that experience for learning.
  • Roll out changes safely with feature flags, design documents, code reviews, tests, observability, and production debugging.
  • Improve engineering quality through clear interfaces, maintainable systems, thoughtful reviews, and strong ownership.
Desired Qualifications
  • Experience with ClickHouse, online analytical processing systems, distributed query engines, or large-scale analytical databases.
  • Experience with RabbitMQ, Temporal, Kafka, Spark, Flink, Ray, Airflow, Dagster, Prefect, or similar queue, workflow, and data systems.
  • Experience with OpenTelemetry, observability products, tracing, logging, or monitoring infrastructure.
  • Experience building systems that call large language model application programming interfaces at scale, including rate-limit management, retries, batching, and cost control.
  • Experience with large language model evaluation, labeling systems, rubric generation, context engineering, reinforcement-learning data pipelines, embedding pipelines, vector search, clustering, or anomaly detection.
  • Experience building developer-facing products, software development kit-backed platforms, or customer-facing infrastructure.

Judgment Labs builds AI reliability infrastructure. Its core offering is an agent behavior monitoring (ABM) platform paired with the open-source Judgeval framework, which traces an AI agent’s entire production workflow—logging LLM calls, tool usage, and memory. It provides automated behavior classification, real-time alerts, and dashboards to track latency and cost, helping teams observe and improve multi-step AI workflows. By allowing rules or human scoring to define correct behavior, it aims to enable safe deployment and continuous improvement of reliable autonomous agents in high-stakes sectors like legal, finance, and enterprise support.

Company Size

11-50

Company Stage

Series A

Total Funding

$32M

Headquarters

San Francisco, California

Founded

2025

Get referred to Judgment Labs

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Lightspeed led both 2026 rounds, signaling conviction after doubling down within six months.
  • The company says the platform already runs at agent-native companies in production.
  • Judgeval open-source adoption can funnel developers into paid monitoring and improvement workflows.

What critics are saying

  • LangSmith, Braintrust, Arize, and Confident AI compress pricing and differentiation in 2026.
  • Public customer proof remains thin; the company disclosed no named enterprise logos by July 2026.
  • If agent observability becomes a LangChain feature, Judgment Labs loses standalone relevance.

What makes Judgment Labs unique

  • May 12, 2026 funding backed Judgeval’s trajectory-level agent monitoring, not just outputs.
  • Shan, Li, and Camyre combine Stanford NLP, TogetherAI research, and Datadog infrastructure experience.
  • The platform targets production failures in legal, finance, and enterprise support workflows.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%
Business Wire
May 13th, 2026
Judgment Labs Closes $32M in Seed and Series A Funding to Build the Continuous Improvement Layer for AI Agents

Today, Judgment Labs, the infrastructure company helping AI-native teams turn production data into continuously improving agents, announced $32 million in co...

FinancialContent
May 12th, 2026
Judgment Labs raises $32M to build continuous improvement layer for AI agents

Judgment Labs, an AI infrastructure startup, has raised $32 million across seed and Series A rounds to build continuous improvement tools for AI agents. Lightspeed Venture Partners led both rounds, with participation from Nova Global, SV Angel, Valor Equity Partners and Dynamic. Founded by three childhood friends — CEO Alex Shan, Chief Scientist Andrew Li and CTO Joseph Camyre, all in their early twenties — Judgment Labs addresses the challenge of evaluating "deep agents" that execute complex tasks rather than simply answering questions. The platform analyses entire agent trajectories to identify failure patterns and helps teams improve their agents using production data. The company's technology is already deployed at agent-native companies. Funding will primarily support hiring AI researchers and engineers in San Francisco.