Contract

Staff Site Reliability Engineer

Doctolib

Doctolib

Healthcare access and practice software

No salary listed

Île-de-France, France + 1 more

More locations: Levallois-Perret, France

Hybrid

Up to 2 remote days per week.

Category
DevOps & Infrastructure (1)
Required Skills
Kotlin
Datadog
Kubernetes
Python
Incident Response
Ruby
Java
OpenTelemetry
TypeScript
AWS
Go
Prometheus
iOS/Swift
Terraform
Observability
React Native
Google Cloud Platform
Requirements
  • Have extensive experience (8+ years) in site reliability engineering, platform engineering, or infrastructure roles within a large-scale, multi-team production environment.
  • Have proven experience with cloud platforms such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure.
  • Have strong experience with containerization and orchestration technologies; Kubernetes is required, including its deployment and scaling strategies ecosystem.
  • Have implemented and operated service-level indicators, service-level objectives, and error budgets in production.
  • Have experience managing on-call rotations and leading incident response in high-stakes environments.
  • Have a strong systems engineering background with fluency in at least one backend programming language, such as Go, Python, or Ruby.
  • Have a proven ability to lead through influence by setting technical direction, driving consensus, and mentoring engineers across teams.
  • Be comfortable balancing long-term architecture work with fast, iterative improvements.
  • Have clear, concise written and verbal communication skills, with the ability to drive alignment in ambiguous environments.
  • Partner with feature teams to accelerate production readiness by providing hands-on guidance on reliability best practices, launch reviews, and operational standards before go-live.
  • Be fluent in English.
Responsibilities
  • Lead large-scale, cross-cutting reliability initiatives across the platform, spanning infrastructure automation, observability, and incident management.
  • Identify and drive improvements to incident detection, response, and postmortem analysis capabilities.
  • Define and evolve service-level objectives, error budgets, and alerting standards across multiple product teams.
  • Participate in the on-call rotation and improve the on-call experience by refining alerting, reducing noise, and ensuring actionable telemetry.
  • Serve as a mentor and technical coach to senior engineers, helping elevate reliability engineering across the company.
  • Influence strategic decisions by providing technical guidance to leadership and representing reliability engineering in architectural reviews and platform discussions.
  • Partner with software engineering teams to embed reliability practices early in the development lifecycle.
Desired Qualifications
  • Have deep expertise in observability tooling and architecture, including logging, tracing, and metrics.
  • Have experience designing and operating high-scale telemetry pipelines and working with developers to improve instrumentation quality.
  • Appreciate working in regulated environments, healthcare, fintech, or similar domains, where data privacy and compliance are part of engineering decisions.
  • Care about reliability enablement, golden paths, runbooks, and shared libraries.

Doctolib is a European healthcare technology company. Its platform supports appointment booking, teleconsultation, patient communication, clinical workflows, and practice administration. Patients use Doctolib to find and manage care, while healthcare professionals and organizations use its software to coordinate services. The company combines consumer applications with subscription software, implementation, security, customer support, and local market operations. Roles span product, engineering, data, design, sales, implementation, customer service, and healthcare operations.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A