Full-Time

Lead Infrastructure & Site Reliability Engineer

Posted on 9/8/2026

Deadline 9/8/27
Crum & Forster Insurance

Crum & Forster Insurance

Specialty insurance and risk solutions

Compensation Overview

$105.4k - $198.1k/yr

+ Equity compensation + Performance-based variable pay

Glastonbury, CT, USA

Remote

Bachelor's

Category
DevOps & Infrastructure (1)
Required Skills
PowerShell
Bash
Microsoft Azure
Python
Grafana
Computer Networking
GraphQL
Vulnerability Analysis
Prometheus
Terraform
Observability
REST APIs
DevOps
Requirements
  • A Bachelor's degree in Computer Science, or a related field, or equivalent experience.
  • At least 8 years of experience in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
  • Deep, hands-on expertise operating production workloads on Microsoft Azure.
  • Proven experience with Infrastructure as Code using Bicep and/or Terraform.
  • Hands-on experience with observability tooling across metrics, logs, and traces, such as Grafana, Prometheus, and Azure Application Insights.
  • Proven experience defining and operating service-level indicators, service-level objectives, and error budgets.
  • Hands-on experience securing Azure cloud environments, including identity and access management, network security, posture management, and vulnerability management.
  • Experience scaling Microsoft Fabric and Microsoft Purview.
  • Experience owning incident management, on-call operations, and blameless post-incident reviews, using tools such as Better Stack, PagerDuty, or comparable platforms.
  • Strong scripting and automation ability using tools such as PowerShell, Python, or Bash, and continuous integration and continuous delivery experience.
  • Experience supporting business continuity and disaster recovery with defined recovery time objectives and recovery point objectives.
  • Experience building internal developer platforms, golden paths, and self-service infrastructure.
  • Cloud cost optimization and FinOps experience.
  • Exposure to AI-assisted operations and modern reliability automation.
  • Strong problem-solving and analytical ability.
  • Strong verbal and written communication skills, including the ability to explain technical and functional issues to technical and non-technical stakeholders.
  • Demonstrated leadership or mentoring of distributed or offshore teams.
Responsibilities
  • Own and evolve TII's Azure infrastructure and reliability roadmap in alignment with the Azure Well-Architected Framework.
  • Co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains with the AVP of Infrastructure.
  • Define standards for compute, network, and platform services that scale with business growth and new distribution channels.
  • Drive cloud cost optimization by balancing performance, resilience, and spend.
  • Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.
  • Own run-time reliability across availability, performance, scalability, and capacity for the platform.
  • Mature and expand the service-level objective practice by defining service-level indicators, refining 99.9% and higher service-level objectives where warranted, and operating error budgets.
  • Lead capacity planning and performance engineering to support platform growth.
  • Drive operational readiness reviews for new services and major releases.
  • Own and mature the observability platform across metrics, logs, and traces by enhancing Grafana, Prometheus, and Azure Application Insights.
  • Implement distributed tracing across GraphQL and REST services to accelerate diagnosis and reduce time to detect and resolve issues.
  • Establish meaningful alerting and telemetry that reduce noise and surface meaningful signals.
  • Build reliability dashboards that provide teams and leadership with visibility into service health.
  • Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
  • Define golden paths and paved-road templates that make reliable, secure, and compliant practices easy to follow.
  • Establish and champion Infrastructure as Code standards using Bicep and Terraform.
  • Improve developer experience and engineering enablement through automation and reusable platform services.
  • Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
  • Establish operational runbooks, self-healing patterns, and proactive reliability practices.
  • Continuously improve deployment safety and rollback capability in partnership with continuous integration and continuous delivery owners.
  • Own the engineering and operational security of the Azure cloud platform, including identity and access management, network security, configuration hardening, and secrets and key management.
  • Manage and improve the cloud security posture using Microsoft Defender for Cloud and Azure Policy, including continuous vulnerability management and remediation.
  • Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
  • Embed DevSecOps and secure-by-design practices into platform and Infrastructure as Code workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
  • Partner with the Security function on policy, governance, and compliance in a regulated insurance environment containing personally identifiable information.
  • Own the major-incident process and incident command on the on-call platform.
  • Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
  • Improve on-call health, escalation paths, mean time to detect, and mean time to resolve.
  • Maintain and enhance business continuity and disaster recovery capabilities, including recovery time objective and recovery point objective targets and periodic testing.
  • Co-lead a small team of infrastructure, cloud, and systems engineers.
  • Uplift reliability and platform capability across engineering by raising standards and building a reliability culture.
  • Mentor engineers and demonstrate leadership behaviors that support growth into a formal infrastructure leadership role.
  • Establish standards, documentation, and ways of working that scale across teams.
Desired Qualifications
  • Microsoft Azure certifications, such as Azure Solutions Architect Expert, Azure DevOps Engineer Expert, or Azure Security Engineer Associate.
  • Experience in insurance, travel, fintech, or software-as-a-service environments, particularly regulated environments.
  • Prior leadership or mentoring of distributed or offshore teams; formal people-leadership experience is a plus.
Crum & Forster Insurance

Crum & Forster Insurance

View

Crum & Forster is a property, casualty, accident and health insurer offering specialty coverage and related risk services. Its businesses underwrite products for organizations, industries and individuals with needs that may not fit standard insurance markets, supported by claims and loss-control teams. The company distributes through brokers, agents and program relationships and is part of Fairfax Financial Holdings. Its operating identity spans specialized underwriting, policy administration and claims rather than functioning as an insurance comparison site or retail agency.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A