Full-Time

Site Reliability Engineer

Axle

Axle

201-500 employees

Translational research informatics and data science

Compensation Overview

$140k - $155k/yr

Frederick, MD, USA

In Person

Category
DevOps & Infrastructure (1)
Required Skills
PowerShell
Chef
Bash
Kubernetes
FedRAMP
MLOps
Microsoft Azure
Python
Grafana
Puppet
GitHub Actions
High Performance Computing (HPC)
NoSQL
R
Node.js
SQL
Java
Data Engineering
Infrastructure as Code (IaC)
Docker
AWS
Jenkins
Terraform
Observability
Ansible
DevOps
Splunk
Google Cloud Platform

Get referred to Axle

See people who can refer or advise you

Requirements
  • At least 6 years of experience in DevOps or Site Reliability Engineering roles using monitoring and observability tools such as Prometheus, Grafana, ELK, or cloud-native equivalents for on-premises and cloud-hosted workloads.
  • At least 4 years of hands-on Linux experience, including Ubuntu, CentOS, or Red Hat operating systems, containers, dependency management, and administration support.
  • At least 4 years of experience automating Infrastructure as Code deployments to Amazon Web Services, Google Cloud Platform, or Microsoft Azure.
  • At least 4 years of experience with continuous integration and continuous delivery and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, or GitHub Actions.
  • Strong scripting skills in Python, Bash, PowerShell, or a similar language.
  • Proficiency using vibe coding and coding assistants to develop scripts, tools, and applications for DevOps and Site Reliability Engineering use cases.
  • Proficiency debugging, troubleshooting, and deploying SQL or NoSQL databases, object storage, web servers, and open-source programming stacks for Node.js, R, Python, .NET Core, or Java.
  • Willingness to learn, adopt, and adapt to emerging technologies and changing project needs.
Responsibilities
  • Design and implement enterprise-grade monitoring and observability frameworks for metrics, logs, and traces across distributed systems using Splunk, Grafana, and OpenTelemetry tools.
  • Establish and manage service-level indicators, service-level objectives, and error budgets to drive reliability improvements.
  • Develop and maintain real-time asset inventory systems across cloud, on-premises, and hybrid environments.
  • Automate workload onboarding and offboarding processes while ensuring standardization and governance.
  • Track system ownership, dependencies, and lifecycle states for operational transparency.
  • Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact.
  • Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments.
  • Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards such as NIST, FISMA, and FedRAMP.
  • Embed ITIL processes, including incident, change, problem, and configuration management, into Site Reliability Engineering workflows.
  • Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards.
  • Develop golden paths and standardized platform templates for consistent workload deployment.
  • Automate provisioning, patching, configuration management, and environment lifecycle management.
  • Use AI/ML coding assistants and vibe coding practices to develop automation scripts, tools, and internal platforms.
  • Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights.
  • Lead adoption of AI-enhanced Site Reliability Engineering practices, including intelligent remediation and predictive operations.
  • Champion DevOps and Site Reliability Engineering practices, including Infrastructure as Code, continuous integration and continuous delivery, observability, and reliability engineering.
  • Build developer-friendly platforms and golden paths that simplify deployments, reduce friction, and improve velocity.
  • Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, inference environments, GPU-enabled workloads, and high-performance computing workloads.
  • Build and manage containerized and orchestrated platforms using Docker and Kubernetes.
  • Support cloud migration, modernization, and platform standardization initiatives.
  • Ensure systems meet security, compliance, backup, and disaster recovery requirements.
  • Promote best practices in DevOps, Site Reliability Engineering, and platform engineering to developer communities.
  • Stay current with technologies including AIOps, MLOps, cloud computing and deployment, Site Reliability Engineering, infrastructure automation, security best practices, and data engineering.
Desired Qualifications
  • Cloud certifications are preferred.
  • Certifications in Grafana, Splunk, Docker, or Kubernetes are preferred but optional.
  • Experience with Java is desired but not mandatory.

Axle Informatics provides specialized informatics solutions for translational research, health informatics, and data science to biomedical research centers and healthcare organizations. Its offerings are customized software and data management platforms that help researchers collect, integrate, analyze, and visualize large research datasets, with end-to-end tools that automate data aggregation and deliver analytics, dashboards, and decision-support features. The company differentiates itself with an integrated, scientifically informed approach that bridges data science with application development, specifically focused on translational research and clinical data work rather than generic software. Axle’s goal is to advance public health by moving biomedical discoveries from the lab to bedside, improving healthcare outcomes.

Company Size

201-500

Company Stage

N/A

Total Funding

N/A

Headquarters

Rockville, Maryland

Founded

2002

Get referred to Axle

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Axle secured a $21 million NCATS support contract starting February 2025.
  • Axle won a 2026 NHLBI systems biology task for single-cell genomics work.
  • Axle’s May 2026 NIH and SK pharmteco partnership expands credibility in rare-disease gene therapy.

What critics are saying

  • NIH concentration makes Axle vulnerable if NCATS and NHLBI rebid work in 2026.
  • Axle filed and lost GAO protests in August 2026, signaling procurement pressure.
  • Federal budget lapses instantly freeze specialized scientific services; Axle’s core contracts depend on appropriated NIH funding.

What makes Axle unique

  • Axle Informatics won NIH NCATS and NHLBI work in 2025-2026.
  • Axle combines biomedical staff scientists, cloud engineering, and translational research program management.
  • Axle’s NIH foothold includes rare-disease gene therapy collaboration with SK pharmteco, May 2026.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Paid Vacation

Paid Holidays

401(k) Company Match

Educational Benefits for Career Growth

Employee Referral Bonus

Flexible Spending Accounts

Company News

TheGWW.com
Jul 25th, 2023
“Octo awarded $64.7M IT Infrastructure Call Order to support NCI’s Cancer research”

Octo, in partnership with Unissant, Axle Informatics, and TRex, has been awarded an IT Infrastructure and Operations Call Order to support the NCI’s OCIO.

GlobeNewswire
Jan 12th, 2023
Digital Pathology Market Worth $1.86 Billion by 2030 -

For instance, in May 2020, Indica Labs (U.S.), a provider of digital pathology solutions, collaborated with information technology consulting companies, Octo (U.S.) and Axle Informatics (U.S.) and the National Institutes of Health (NIH) (U.S.), to develop an online collection of high-resolution histopathology images of tissues from COVID-19 patients using Indica’s HALO Link platform.