Full-Time

Lead Site Reliability Engineer

CardWorks Servicing

CardWorks Servicing

Compensation Overview

$146k - $162.3k/yr

Orlando, FL, USA + 3 more

More locations: South Jordan, UT, USA | Woodbury, NY, USA | Pittsburgh, PA, USA

Hybrid

Hybrid or fully remote work may be considered based on the hiring manager's decision and role priorities.

Master's

Category
DevOps & Infrastructure (1)
Required Skills
LLM
Kubernetes
Microsoft Azure
GitHub Actions
Machine Learning
Infrastructure as Code (IaC)
Docker
AWS
Jenkins
Terraform
Observability
VMWare
Ansible
DevOps
Requirements
  • Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.
  • Demonstrated experience standing up or significantly maturing an SRE practice, including the operating model, SRE service engagement, production readiness, incident and postmortem program, and reliability roadmap.
  • Hands-on experience applying artificial intelligence and machine learning to operations or generative artificial intelligence in production support workflows, with a focus on measurable outcomes such as mean time to detect, mean time to resolve, alert fatigue reduction, and change failure rate, together with responsible-use controls.
  • Proven ability to establish Service Level Indicators and Service Level Objectives in production environments, including hands-on definition and implementation.
  • Demonstrated background in production incident response, leading resolution efforts, conducting blameless post-incident reviews, and implementing actionable remediation strategies.
  • Strong observability and telemetry expertise in designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.
  • Infrastructure engineering experience with strong Infrastructure as Code skills using Terraform and Ansible.
  • Thorough understanding and practical experience in continuous integration and continuous delivery pipeline design, optimization, and troubleshooting using modern tooling and platforms such as Azure DevOps, GitHub Actions, Jenkins, or GitLab CI, with an emphasis on speed, reliability, and security.
  • Practical knowledge of containerization and platform modernization, including architecting and operating containerized workloads with Docker, VMware, and Kubernetes or comparable orchestration platforms, to modernize legacy applications and improve fault tolerance.
  • Knowledge of emerging reliability practices, including Service Level Objective automation platforms, AIOps, or predictive operations, to advance proactive reliability management.
  • Master’s degree in computer science, engineering, or equivalent practical experience designing and operating production systems at scale.
  • At least 7 years of experience in Site Reliability Engineering.
Responsibilities
  • Establish the SRE operating model, including service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning, and ensure adoption across teams.
  • Identify, pilot, and operationalize artificial-intelligence-enabled reliability use cases such as alert noise reduction, incident summarization, correlation and root-cause hypothesis generation, runbook assistance, and human-approved auto-remediation with appropriate guardrails.
  • Define, implement, and operationalize reliability metrics by establishing and managing Service Level Indicators, Service Level Objectives, and error budgets to quantify and continuously improve service reliability and support engineering and business decisions.
  • Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake and prioritization process aligned to business criticality.
  • Define and enforce error budget policies, including escalation paths and release-risk decisions, in partnership with Product and Engineering, using Service Level Objective attainment to guide trade-offs between feature velocity and reliability.
  • Establish and maintain centralized paved-road reliability standards and assets, including instrumentation conventions, golden signals, alerting standards, runbook templates, and Service Level Objective dashboards, that product teams can adopt with minimal friction.
  • Design the on-call and escalation model for a centralized SRE team, including an SRE overlay for major incidents, defined handoffs with service owners, and clear ownership boundaries.
  • Design and engineer automation and observability solutions by developing tooling, dashboards, and systems to reduce operational toil, enhance system visibility, and accelerate delivery.
  • Participate in incident and problem management as incident coordinator for high-severity events, drive cross-functional responses, conduct blameless root-cause analysis, run post-incident reviews with clear owners and due dates, and ensure remedial actions improve reliability.
  • Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.
  • Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements and ensure secure, auditable, and well-documented practices.
  • Collaborate with end users, product management, development, architecture, and IT Operations teams to embed reliability principles throughout the software development lifecycle, including service onboarding, reliability reviews, and shared Service Level Objective ownership.
  • Promote reliability throughout all phases of development, advocate for continuous improvement, and communicate key metrics and potential customer impact to stakeholders.
  • Train, mentor, and upskill engineering teams in SRE practices, support junior team members, foster shared ownership and accountability for reliability, and influence teams without direct authority through standards, data, and executive-aligned priorities.
  • Remain current on SRE trends and best practices, including observability, AIOps, and Service Level Objective management, and implement these methodologies to support business outcomes.
  • Evaluate artificial intelligence tools for reliability with security, privacy, and compliance guardrails, including data handling, prompt and content controls, and auditability, and measure their impact.
  • Participate in on-call rotations and operational support for SRE-supported systems and products.
Desired Qualifications
  • AWS Professional certification, Terraform certification, Ansible certification, Azure DevOps certification, Octopus Deploy certification, or other automation-focused credentials demonstrating continuous technical development.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A