Full-Time

AI Infrastructure Engineer

Cluster Engineer

Updated on 9/4/2026

STN

STN

51-200 employees

Vendor-agnostic IT infrastructure and GPU services

No salary listed

Remote in USA + 1 more

More locations: San Francisco, CA, USA

Remote

Category
DevOps & Infrastructure (1)
Required Skills
Bash
Kubernetes
Microsoft Azure
Python
Grafana
CUDA
PyTorch
AWS
Prometheus
Terraform
Ansible
Linux/Unix
Google Cloud Platform

Get referred to STN

See people who can refer or advise you

Requirements
  • At least 7 years of experience designing or operating large-scale Linux infrastructure.
  • At least 5 years of experience supporting production GPU clusters for artificial intelligence or high-performance computing workloads.
  • Demonstrated experience building multi-node GPU training environments from the ground up.
  • Deep expertise with distributed PyTorch training.
  • Extensive experience troubleshooting and optimizing NCCL communications.
  • Strong understanding of distributed artificial intelligence communication patterns, including AllReduce, ReduceScatter, AllGather, Broadcast, and point-to-point communications.
  • Experience benchmarking distributed training using nccl-tests, NVIDIA DCGM, and Nsight Systems.
  • Strong understanding of GPU memory management, including KV Cache, activation checkpointing, tensor parallelism, pipeline parallelism, and data parallelism.
  • Experience optimizing large language model inference throughput, including tokens-per-second optimization, batch sizing, continuous batching, KV cache tuning, and memory bandwidth optimization.
  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
  • Expert-level Linux systems administration skills.
  • Experience with the Slurm workload manager.
  • Experience using Pyxis and Enroot for containerized GPU workloads.
  • Strong scripting skills using Python and Bash.
  • Experience with InfiniBand, RoCE v2, RDMA, GPUDirect RDMA, GPUDirect Storage, UCX, MPI, network topology optimization, congestion control, quality of service, ECN/PFC, and high-speed Ethernet.
  • Experience designing or tuning storage for AI workloads, including parallel file systems, distributed storage, object storage, NVMe, checkpoint optimization, dataset staging, storage bandwidth optimization, and metadata performance.
  • Experience with multi-node scaling analysis, GPU utilization optimization, communication/computation overlap, NUMA optimization, CPU affinity, PCIe topology, GPU topology including NVLink/NVSwitch, memory bandwidth analysis, and end-to-end performance profiling.
Responsibilities
  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
  • Optimize inference clusters for token-generation throughput, low latency, and high GPU utilization.
  • Build and support production AI infrastructure running hundreds to thousands of GPUs.
  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and quality of service.
  • Configure and tune distributed AI software stacks including PyTorch, NCCL, CUDA, UCX, MPI, Slurm, and Pyxis/Enroot.
  • Optimize GPU scheduling and resource allocation for training and inference environments.
  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
  • Identify performance regressions and troubleshoot distributed training issues at scale.
  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
  • Work closely with machine learning engineers to improve training scalability and inference efficiency.
  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.
  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.
Desired Qualifications
  • Experience deploying AI workloads on Kubernetes.
  • Experience with NVIDIA GPU Operator.
  • Experience with Kubernetes batch scheduling tools such as Volcano, Kueue, and Run:ai.
  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
  • Familiarity with MLPerf benchmarking.
  • Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
  • Experience automating infrastructure using Ansible, Terraform, or similar tools.
  • Experience working in cloud GPU environments including AWS, Azure, and GCP in addition to bare metal.
  • Experience with Triton and TensorRT-LLM.

STN provides IT infrastructure services for enterprise IT and AI workloads. It designs, runs, and protects technology foundations through end-to-end services like Technology Consulting, Managed IT Services, and Technology Procurement via its VAR. Its GPU One platform offers private GPU-as-a-Service for scalable AI development, and STN One handles managed security, incident response, and risk/compliance. By being vendor-agnostic, STN crafts customized, secure, high-performance environments for sectors such as financial services, healthcare, government, and engineering.

Company Size

51-200

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

2016

Get referred to STN

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • STN listed on CRN Solution Provider 500 in 2026 for the fourth straight year.
  • STN launched early access for NVIDIA B300 on September 4, 2026, signaling demand.
  • STN states active GPU One locations in Silicon Valley, Los Angeles, and Chicago.

What critics are saying

  • STN's 2026 vision spans robotics, autonomous data centers, and global AI utility, diluting focus.
  • GPU One still relies on colocation providers; Missouri, South Carolina, and Texas remain under development.
  • STN's growth narrative depends on NVIDIA supply access and partner status, both concentrated dependencies.

What makes STN unique

  • STN runs GPU One on validated NVIDIA H200, B200, B300, and GB300 architecture.
  • STN pairs managed IT, cybersecurity, procurement, and GPU cloud under one operating model.
  • STN markets vendor-agnostic roadmapping for regulated buyers in healthcare, finance, and government.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote Work Options

Hybrid Work Options