Full-Time

HPC Infrastructure Site Reliability Engineer

Updated on 9/4/2026

Radiant

Radiant

51-200 employees

Operates sovereign AI infrastructure and compute

No salary listed

London, UK + 2 more

More locations: Gloucestershire, UK | United Kingdom

Hybrid

Hybrid work is indicated for Gloucestershire and London.

Bachelor's, Master's

Category
DevOps & Infrastructure (1)
Required Skills
TCP/IP
Graphics Processing Unit (GPU)
Bash
Kubernetes
Python
Grafana
Distributed Systems
High Performance Computing (HPC)
CUDA
Computer Networking
Infrastructure as Code (IaC)
Prometheus
Observability
Ansible
Linux/Unix

Get referred to Radiant

See people who can refer or advise you

Requirements
  • At least 8 years of experience in Site Reliability Engineering, Infrastructure Engineering, or similar roles in large-scale distributed production environments operating a 24/7 support model.
  • At least 2–3 years of recent experience in high-performance computing and/or artificial intelligence infrastructure, including GPU-based compute environments at scale.
  • Strong Linux expertise, preferably Ubuntu, including deep systems administration and production troubleshooting.
  • Proven experience tuning compute-system performance, including kernel, BIOS/firmware, and storage subsystem optimization.
  • Strong hands-on experience with bare-metal infrastructure and out-of-band management tooling such as IPMI, iLO, iDRAC, Redfish, or equivalent.
  • Solid networking fundamentals including TCP/IP, DNS, DHCP, VLANs, routing, and switching, with exposure to high-performance networking environments.
  • Exposure to NVIDIA GPU ecosystems, including CUDA-based workloads, GPU-accelerated compute environments, and the NVIDIA AI reference architecture.
  • Familiarity with high-performance networking technologies such as InfiniBand and RoCE.
  • Strong experience with infrastructure automation and scripting, such as Bash, Python, Ansible, or similar infrastructure-as-code and tooling approaches.
  • Understanding of observability principles and practical use of monitoring and telemetry systems such as Prometheus, Grafana, or equivalents.
  • Understanding of workload schedulers and running workloads across multiple systems in parallel.
  • Practical experience with at least one parallel storage platform.
  • Experience working in ITIL-aligned environments, including Incident, Major Incident, Problem, and Change Management.
  • Strong troubleshooting skills in high-pressure operational environments, with a track record of incident ownership and resolution.
  • Strong communication skills with the ability to work across engineering teams and interface with non-technical stakeholders.
  • Ability and willingness to collaborate closely with Platform Site Reliability Engineering teams, including exposure to and learning of Kubernetes-based orchestration environments.
  • Experience contributing to or influencing infrastructure design, reliability improvements, or operational best practices.
Responsibilities
  • Operate and improve high-density artificial intelligence and high-performance computing infrastructure in a 24/7 production environment.
  • Participate in a 24x7x365 on-call rotation, supporting mission-critical systems and incident response.
  • Troubleshoot complex issues across compute, networking, storage, and orchestration layers in GPU-accelerated environments.
  • Lead performance evaluation, testing, and operational acceptance of new high-performance computing infrastructure before production release.
  • Drive continuous service improvement by reducing toil through automation, tooling, and process refinement.
  • Build and maintain infrastructure automation and tooling using infrastructure as code and scripting to improve reliability and operational efficiency.
  • Optimize Linux systems for performance, including kernel, BIOS/firmware, and storage tuning for high-performance computing workloads.
  • Configure and operate bare-metal infrastructure using IPMI, iLO, iDRAC, Redfish, and related tooling.
  • Partner with infrastructure tooling and observability teams to improve telemetry, alerting, and system visibility at scale.
  • Own ITIL-aligned processes across Incident, Major Incident, Problem, and Change Management, ensuring strong execution and continuous improvement.
  • Lead root cause analysis and ensure corrective actions are implemented and automated where possible.
  • Help design and deliver future high-performance computing cluster and site builds, shaping global consistency and operational standards.
  • Collaborate with Platform Engineering, Network Engineering, Infrastructure Tooling, and Data Centre Operations to improve reliability and deployment quality.
  • Feed operational insight back into infrastructure design to influence next-generation high-performance computing platform evolution.
  • Mentor engineers and act as a technical authority for operational best practices across teams.
  • Communicate technical and non-technical issues as actionable outcomes for stakeholders.
  • Uphold a culture of doing, documenting, and automating.
Desired Qualifications
  • Deep experience with high-performance computing workloads and GPU-accelerated infrastructure at scale.
  • Experience with InfiniBand, RoCE, or other high-performance computing-grade networking fabrics in production environments.
  • Experience with high-performance computing benchmarking, validation, or performance testing, including tools such as LINPACK, FIO, NCCL, or ibdiagnet.
  • Exposure to large-scale multi-site or global infrastructure deployments.
  • A Bachelor or Masters level degree in Computer Science, Engineering, or a related field, or equivalent on-the-job experience.
  • LPIC certifications.
  • ITIL Foundation level qualification or equivalent experience.

Radiant builds and operates utility-scale AI infrastructure to deliver sovereign AI cloud compute through owned data centers, long-term capital, and proprietary software. It runs AI factories that provide scalable, resilient compute for governments, enterprises, and telecom providers, enabling rapid deployment of AI workloads at scale. After merging with Ori Industries, Radiant integrates distributed AI software with its hardware to form a global AI utility that lowers energy costs and reduces dependence on centralized clouds. Its goal is to create a secure, scalable AI platform that serves distributed locations worldwide, balancing land, capital, and software to support large-scale AI applications.

Company Size

51-200

Company Stage

N/A

Total Funding

N/A

Headquarters

London, United Kingdom

Founded

2026

Get referred to Radiant

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Feb. 24, 2026 launch created a $1.3 billion platform with rolled-in Ori investors.
  • Radiant claims direct access to Brookfield's $100 billion AI infrastructure program.
  • Ori Global AI Cloud keeps on-demand bare metal, GPU instances, and fine-tuning offerings live.

What critics are saying

  • Radiant depends on Brookfield funding and contract wins to justify its capital-intensive model.
  • NVIDIA chip supply concentration and rapid obsolescence can strand Radiant's data-center investments.
  • Execution failure on sovereign contracts or deployment delays destroys the utility-scale thesis entirely.

What makes Radiant unique

  • Feb. 24, 2026 merger fused Ori software with Brookfield-powered infrastructure.
  • Radiant pairs powered land, long-term capital, and sovereign-cloud software under one operator.
  • NVIDIA DSX, Blackwell, and GB200 NVL72 make Radiant a purpose-built AI factory platform.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Paid Vacation

Parental Leave

Gym Membership

Commuter Benefits

Company Equity

Professional Development Budget

Growth & Insights and Company News

Headcount

6 month growth

-11%

1 year growth

-11%

2 year growth

-11%
Tech.eu
Feb 24th, 2026
Radiant and Ori merge to deliver sovereign AI cloud at utility scale

The deal unites distributed AI cloud software with powered land and long-term capital, positioning Radiant as a vertically integrated AI infrastructure player.