Full-Time

Data Center Operations and Maintenance Engineer

Designworks Talent

Designworks Talent

No salary listed

No H1B Sponsorship

Bellevue, WA, USA

Hybrid

Approximately three days per week in the office.

Category
DevOps & Infrastructure (1)
Required Skills
Datadog
Graphics Processing Unit (GPU)
Grafana
Incident Response
Computer Networking
Prometheus
Requirements
  • Experience in a data center operations, site reliability, or infrastructure operations role, ideally supporting GPU or large-scale compute environments.
  • Strong incident response and troubleshooting skills across hardware, networking, and systems layers.
  • Comfort working in an early-stage environment where processes and tooling are still being established.
  • Experience in multiple technical environments is a plus.
Responsibilities
  • Monitor live data center and GPU infrastructure for performance, capacity, and reliability issues.
  • Lead or support incident response for production issues, driving toward fast, effective resolution.
  • Collaborate with engineering teams to ensure monitoring, alerting, and operational tooling are first class and identify issues before they affect customers.
  • Partner with hardware, networking, orchestration, and infrastructure teams to resolve root causes and prevent recurrence.
  • Contribute to runbooks, on-call practices, and operational maturity as the platform scales.
Desired Qualifications
  • Experience with on-call and incident-management tooling such as PagerDuty and Opsgenie and monitoring stacks such as Prometheus, Grafana, and Datadog.
  • Background supporting GPU cluster or data center operations post-launch.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A