Full-Time

Senior Site Reliability Engineer

Platform

Updated on 8/21/2026

Order.co

Order.co

201-500 employees

Digital platform streamlining enterprise procurement lifecycle

Compensation Overview

$175k - $200k/yr

+ Bonus + Equity

Remote in USA

Remote

Category
DevOps & Infrastructure (1)
Required Skills
Datadog
Claude
Bash
Kubernetes
Python
Incident Response
Distributed Systems
Data Structures & Algorithms
Ruby
Ruby on Rails
Computer Networking
OpenTelemetry
Infrastructure as Code (IaC)
CloudFormation
Microservices
AWS
Go
Terraform
REST APIs
DevOps
Linux/Unix

Get referred to Order.co

See people who can refer or advise you

Requirements
  • Demonstrate accountability by owning outcomes rather than only tasks.
  • Measure success by shipping working software and take responsibility for correctness in code affecting customer balances.
  • Design and build software incrementally without requiring a complete specification.
  • Apply AI-driven solutions pragmatically and evaluate AI output critically.
  • Have a strong foundation in computer science fundamentals, including data structures, algorithms, and system design.
  • Have familiarity with production-grade applications and services using Ruby and Ruby on Rails.
  • Have deep expertise in Linux systems administration and production troubleshooting.
  • Have strong experience operating cloud infrastructure at scale, particularly in Amazon Web Services environments.
  • Have experience with Kubernetes, container orchestration, and cloud-native infrastructure patterns.
  • Be proficient with infrastructure as code tools such as Terraform or CloudFormation.
  • Have expertise designing and operating continuous integration and continuous delivery pipelines and deployment automation systems.
  • Have deep understanding of observability tooling including Datadog and OpenTelemetry or similar platforms.
  • Have strong knowledge of distributed-systems reliability patterns, including redundancy, failover, autoscaling, rate limiting, and graceful degradation.
  • Have experience building automation and operational tooling using Python, Go, Bash, or Ruby.
  • Have strong understanding of networking fundamentals, including DNS, load balancing, TLS, VPNs, firewalls, and service discovery.
  • Have hands-on experience with incident response, root-cause analysis, and production operations in high-availability environments.
  • Be familiar with site reliability engineering methodologies, including SLOs, SLIs, error budgets, capacity planning, and operational maturity modeling.
  • Have experience implementing secure infrastructure and cloud-security best practices, including IAM, secrets management, and vulnerability remediation.
  • Have proven ability to design scalable, resilient, and maintainable platform systems and APIs.
  • Have experience supporting distributed microservices architectures and event-driven systems.
  • Have strong understanding of operational excellence principles, including automation-first engineering and toil reduction.
  • Have experience using AI-assisted engineering tools such as Claude or GitHub Copilot while applying sound operational and engineering judgment.
  • Demonstrate debugging and systems-thinking skills across infrastructure, networking, application, and platform layers.
Responsibilities
  • Design, build, and operate highly available, scalable, and fault-tolerant infrastructure and platform services.
  • Own reliability, availability, latency, and operational excellence for critical production systems and services.
  • Define and maintain service-level objectives, service-level indicators, and error budgets across platform systems.
  • Lead incident response efforts for complex production outages and drive root-cause analysis and long-term remediation actions.
  • Build resilient systems that handle failures, traffic spikes, dependency degradation, and regional outages.
  • Continuously improve system reliability through automation, observability, performance tuning, and capacity planning.
  • Develop infrastructure automation and self-service tooling to reduce operational toil and improve engineering velocity.
  • Build and maintain continuous integration and continuous delivery pipelines, deployment automation, and release-engineering workflows.
  • Implement infrastructure as code practices using Terraform, CloudFormation, and container orchestration.
  • Build reliable internal platforms, operational tooling, and standardized deployment patterns to improve developer experience.
  • Drive adoption of GitOps, immutable infrastructure, and automated remediation patterns.
  • Design and maintain monitoring, logging, tracing, and alerting systems for distributed services.
  • Establish actionable alerting standards that reduce noise and improve incident detection and response times.
  • Analyze production trends, system bottlenecks, and failure patterns to proactively prevent incidents.
  • Lead operational readiness reviews, disaster-recovery planning, and game-day exercises.
  • Improve mean time to detect and mean time to recovery through tooling, automation, and process refinement.
  • Participate in architecture and infrastructure design reviews.
  • Propose scalable and reliable platform designs covering multi-region deployment, redundancy, failover, and security considerations.
  • Evaluate trade-offs among reliability, scalability, operational complexity, and engineering velocity.
  • Identify systemic risks and operational gaps before they become production incidents.
  • Partner with engineering teams to ensure services are designed for operability, observability, and resilience from the outset.
  • Implement and maintain secure cloud networking, secrets management, IAM policies, and infrastructure-hardening standards.
  • Partner with Security and Compliance teams to ensure systems meet organizational and regulatory requirements.
  • Drive operational best practices for vulnerability management, patching, and production access controls.
  • Scope and estimate infrastructure and reliability initiatives accurately.
  • Coordinate production rollouts, maintenance events, and reliability improvements across teams.
  • Communicate operational risks, dependencies, and incident impacts to technical and non-technical stakeholders.
  • Collaborate with Software Engineering, Security, Product, and Operations teams to improve platform reliability and scalability.
  • Serve as a trusted escalation point during critical production incidents.
  • Mentor junior and mid-level engineers on reliability engineering, operational excellence, and infrastructure best practices.
  • Raise the engineering organization's operational maturity through documentation, reviews, and technical guidance.
  • Drive improvements in team standards around observability, incident management, automation, and infrastructure design.
  • Influence technical decisions through operational expertise and engineering judgment.
Desired Qualifications
  • Prior hands-on use of AI-assisted development tools.

Order.co provides a digital platform that centralizes and streamlines procurement for enterprises, helping organizations manage purchasing across the full procurement lifecycle. The product integrates requests, approvals, supplier catalogs, purchase orders, spend tracking, and reporting, typically syncing with ERP or finance systems so users route approvals, order goods or services, receive invoices, and analyze spend in one place. It differentiates itself by offering end-to-end control and visibility in a single system for requisitions, supplier management, and analytics, which enforces company policies and reduces shadow procurement. Goal: help enterprises save time, cut unnecessary spend, and improve transparency and governance over purchasing processes.

Company Size

201-500

Company Stage

Series B

Total Funding

$85.6M

Headquarters

New York City, New York

Founded

2014

Get referred to Order.co

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • March 2026 Hackett Group recognition strengthens credibility with procurement buyers.
  • June 2026 Workday integration and upcoming built-on Workday app expand enterprise reach.
  • Recent 2026 funding and $332 million valuation support hiring and product investment.

What critics are saying

  • Generic AI procurement rivals will copy Command Center features and compress pricing by 2027.
  • Workday platform dependence concentrates distribution risk if Workday changes partner priorities.
  • If savings and compliance metrics stall, enterprises will replace Order.co with native ERP tools.

What makes Order.co unique

  • Order.co’s AI Command Center beta, launched November 2025, automates finance and procurement tasks.
  • Its Workday innovation partner integration, updated June 2026, embeds buying inside Workday.
  • The multi-supplier checkout removes punchout sprawl, a strong workflow advantage for enterprises.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Wellness Program

401(k) Company Match

Remote Work Options

Flexible Work Hours

Parental Leave

Stock Options

Company Equity

Hybrid Work Options

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

-6%
Spend Matters
May 27th, 2025
Aligning Finance And Procurement – Phase 4: Leveraging Technology To Bridge The Finance–Procurement Gap

Finance and procurement play critical roles in shaping an organization’s financial health and operational efficiency. This series shapes a five-phase approach to aligning finance and procurement.Using technology to bridge the Finance-Procurement gapOnce a well-defined technology adoption strategy is established in the ‘collaboration’ phase, organizations can move to the next step: leveraging technology to bridge the finance-procurement gap.Technology is a powerful enabler of finance-procurement collaboration, bridging gaps in spend visibility, cost control and data integration. Without the right tools, procurement’s contributions to financial strategy can remain disconnected from budgeting, forecasting and risk management.By leveraging automation, AI-driven analytics and integrated source-to-pay (S2P) solutions, organizations can streamline workflows, improve financial accuracy and enhance procurement’s role in decision making. The following sections explore how technology enables better finance-procurement integration, the key considerations for selecting the right solutions and how digital transformation drives efficiency, compliance and long-term financial impact.The importance of integrating procurement and finance systemsA major challenge for organizations is siloed procurement and financial systems which separately manage critical processes, such as spend tracking, budget planning and supplier risk assessment. Without integration, finance teams struggle to track procurement-driven savings, while procurement lacks visibility into financial planning and liquidity management.To resolve this, companies need to do the following:1. Ensure procurement data is linked to financial dashboards

PR Newswire
Jan 25th, 2022
Negotiatus Raises $30M in Series B Funding, Rebrands to Order

/PRNewswire/ -- Negotiatus, a guided B2B marketplace simplifying the buying process for businesses, today announced it has raised an additional $30M in funding...

Order.co
Jan 25th, 2022
Stage2 invests into Negotiatus in $30M

Negotiatus is also thrilled to announce a $30M Series B financing led by Stage 2 Capital via a direct investment from MIT’s Endowment, a $30B+ fund with an impressive track record investing in transformative companies.

Negotiatus
Oct 19th, 2021
Negotiatus hires Corey Beale as Chief Revenue Officer

Negotiatus is pleased to welcome Corey Beale as Negotiatus’ new Chief Revenue Officer, effective immediately.