Full-Time

Machine Learning Operations Engineer

Updated on 8/1/2026

White Circle

White Circle

11-50 employees

AI production safety and governance platform

Compensation Overview

$100k - $200k/yr

London, UK + 1 more

More locations: Île-de-France, France

Hybrid

Paris-based roles offer relocation support; London-based roles do not.

Category
DevOps & Infrastructure (2)
,
Required Skills
LLM
Datadog
Kubernetes
MLOps
Rust
CUDA
Machine Learning
OpenTelemetry
Prometheus
Terraform
Observability
REST APIs
DevOps

Get referred to White Circle

See people who can refer or advise you

Requirements
  • Experience with an inference serving engine such as SGLang, vLLM, Dynamo, or TensorRT-LLM, and a working understanding of the request lifecycle through the gateway, router, frontend, worker, queue, and model engine.
  • Solid Kubernetes GPU experience, including the NVIDIA device plugin, GPU scheduling, resource requests and limits, node affinity, taints, tolerations, and node pools.
  • Understanding of multi-node communication libraries and kernels, the CUDA runtime, and container runtime compatibility, with the ability to debug across those layers.
  • Ability to design and implement continuous integration and continuous delivery for model serving, including image and configuration versioning, smoke tests, quality regression tests against benchmarks, latency and throughput gates, canary rollout, and rollback.
  • Ability to define dashboards and alerts using observability metrics such as p50, p95, and p99 latency, time to first token, time per output token, queue depth, GPU utilization and memory, error, timeout, and out-of-memory rates, fallback rate, route distribution, canary versus baseline, and cost per successful request.
  • Production debugging across the whole stack, from Rust to Kubernetes configurations.
  • Ability to clearly communicate engineering tradeoffs.
Responsibilities
  • Integrate new text and multimodal models into serving paths and verify that they behave correctly under production-like traffic.
  • Build and maintain rollout pipelines for frequent model releases.
  • Create smoke, quality, and performance gates for model promotion.
  • Operate local and cluster GPU deployments on Kubernetes.
  • Build dashboards for latency, throughput, queue depth, GPU usage, fallback rate, and quality drift.
  • Run A/B and canary rollouts for model, prompt, routing, and serving configuration changes.
  • Debug production issues across model configuration, tokenizer, serving API, router, queue, Kubernetes, GPU runtime, and continuous integration jobs.
  • Optimize serving cost and reliability across mixed GPU capacity.
Desired Qualifications
  • Rust backend experience.
  • Experience with NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA.
  • Experience with ClickStack or Datadog.
  • Experience using Terraform for GPU infrastructure.
  • Experience with DCGM exporter, Prometheus, or OpenTelemetry.
  • Experience with a high model rollout cadence of 2–3 releases per week.

White Circle provides a control layer platform for AI systems in production that sits in front of models via a single API to test, protect, observe, and improve AI behavior. It uses proprietary models to monitor inputs and outputs in real time against client policies, detecting unsafe inputs, jailbreaks, prompt injections, harmful content, hallucinations, model drift, and malicious user activity, and can block or filter as needed. The platform emphasizes governance and observability for enterprise use, designed for both technical and non-technical teams, with security certifications (SOC 2 Type I & II, HIPAA). Its goal is to help finance and healthcare companies reduce risk while improving reliability and compliance across production AI deployments.

Company Size

11-50

Company Stage

Seed

Total Funding

$11M

Headquarters

Paris, France

Founded

2025

Get referred to White Circle

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • $11 million seed from OpenAI, Anthropic, DeepMind, and Hugging Face leaders.
  • SOC 2 Type I and II plus HIPAA compliance opens regulated buyers.
  • One billion API requests and two large digital banks show enterprise traction.

What critics are saying

  • Foundation models can absorb guardrail features and reduce White Circle’s standalone value.
  • Structured-output deployments weaken refusals, creating liability in White Circle’s core workflow.
  • A visible bank or healthcare failure would damage trust and trigger procurement freezes.

What makes White Circle unique

  • Single-API control layer sits between users and models in real time.
  • Proprietary monitoring models detect jailbreaks, hallucinations, drift, and malicious behavior.
  • Founded after Denis Shilov’s universal jailbreak exposed major model safety gaps.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Paid Time Off

Hybrid Work Options

Relocation Assistance

Health Insurance

Company Equity

Team Social Events

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%