Full-Time

HPC Support Engineer

Lambda

Lambda

501-1,000 employees

Cloud-based GPU services for AI training

Compensation Overview

$122k - $162k/yr

+ Equity compensation + Wellness stipend + Commuter stipend + 401(k) match

Remote in USA

Remote

Category
DevOps & Infrastructure (1)
Required Skills
TCP/IP
Datadog
Kubernetes
Grafana
Distributed Systems
High Performance Computing (HPC)
CUDA
Machine Learning
Computer Networking
Docker
Prometheus
Terraform
Ansible
DevOps
Linux/Unix

Get referred to Lambda

See people who can refer or advise you

Requirements
  • At least 3 years of hands-on high-performance computing experience in an administration, support, or engineering role.
  • Strong Linux system administration experience.
  • Experience administering Linux clusters in high-performance computing environments, preferably with Kubernetes or Slurm for cluster orchestration.
  • Strong coding ability and continuous integration and continuous delivery experience, including use of artificial-intelligence-assisted tools.
  • Proficiency with Prometheus, Grafana, or Datadog for monitoring and logging.
  • Strong log analysis, kernel-level debugging, and performance profiling skills.
  • Experience with CUDA, NCCL, NVLink, or GPUDirect RDMA.
  • Experience with high-throughput networking technologies such as InfiniBand or RoCE.
  • Knowledge of distributed artificial intelligence, machine learning, or high-performance computing workloads.
  • Knowledge of TCP/IP, VPNs, and firewalls in cloud environments.
  • Ability to work independently and mentor junior support engineers.
Responsibilities
  • Serve as a senior technical escalation point for infrastructure and platform issues, troubleshooting through the hardware, driver, or kernel level when needed.
  • Distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration to resolve issues correctly.
  • Identify and address process, tooling, and documentation gaps.
  • Use artificial-intelligence tools to build scripts, automations, or small internal tools that address operational gaps.
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure.
  • Document solutions and contribute to evolving support procedures.
  • Collaborate with engineering teams to turn recurring customer pain points into permanent fixes.
  • Take escalations from peers while training and mentoring them.
  • Participate in a rotating on-call schedule and own major incidents and major customer issues.
  • Support high-volume deployments and other operational needs.
Desired Qualifications
  • Experience with virtualization and container technologies such as Docker and Kubernetes.
  • Experience with neoclouds or GPU cloud providers.
  • Flexible availability for shifts outside normal working hours or on weekends.
  • Experience with high-performance storage systems.
  • Familiarity with infrastructure-as-code tools such as Terraform or Ansible.
  • Experience with Nvidia GPUs and InfiniBand.

Lambda Labs provides cloud-based GPU services for AI training and inference. Its AI Developer Cloud runs on NVIDIA GH200 Grace Hopper hardware to train large language models and generative AI, offering on-demand and reserved GPUs billed by the hour (for example, $1.99/hour for H100). It differentiates itself through competitive pricing, high availability, and an integrated ML stack with Lambda Stack for easy installation of PyTorch, TensorFlow, CUDA, cuDNN, and NVIDIA drivers, plus Lambda Echelon for owning infrastructure with hosting and support. Its goal is to make scalable AI development and deployment affordable by providing flexible GPU access, reliable hosting, and streamlined software deployment for teams working with large models.

Company Size

501-1,000

Company Stage

Debt Financing

Total Funding

$4.1B

Headquarters

San Jose, California

Founded

2012

Get referred to Lambda

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • On May 7, 2026, Lambda closed a $1 billion J.P. Morgan credit facility.
  • November 2025's $1.5 billion Series E funds chip buys and data-center builds.
  • May 2026 CEO Michel Combes and Chairman John Donovan bring public-company execution experience.

What critics are saying

  • Lambda depends on Nvidia supply; Blackwell or Rubin delays hit 2026 revenue.
  • The 3GW-by-2030 plan requires flawless data-center execution; slippage destroys margins.
  • Microsoft, CoreWeave, and AWS squeeze pricing; one client cancellation exposes concentration.

What makes Lambda unique

  • In May 2026, Lambda launched bare-metal NVIDIA Vera Rubin NVL72 instances.
  • Lambda's March 2026 GB300 photonics stack targets frontier training and inference workloads.
  • On May 20, 2026, Lambda won Hudson River Trading for 1,000 Blackwell systems.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

401(k) Company Match

Unlimited Paid Time Off

Wellness Program

Commuter Benefits

Growth & Insights and Company News

Headcount

6 month growth

1%

1 year growth

2%

2 year growth

2%
AD HOC NEWS
Aug 11th, 2026
Nvidia arranges $500B AI infrastructure credit, raising Ouroboros risk concerns over chipmaker-lender overlap

Nvidia is arranging over $500 billion in AI infrastructure credit alongside Apollo Global Management, Blackstone, BlackRock, Brookfield Asset Management, Goldman Sachs and KKR, raising concerns about circular financing. The chipmaker may secure up to $125 billion of the total itself. A recent $917 million credit facility with Lambda demonstrates how the system works: customers lacking liquidity can access capital arranged through the consortium to purchase Nvidia hardware. CEO Jensen Huang aims to transform the company's graphics chips into an "investable asset class". Critics warn of an "Ouroboros dilemma" — when manufacturers arrange loans for customers to buy their own products, creating a closed loop vulnerable to collapse if actual AI service demand lags infrastructure buildout. Nvidia shares closed at €188.50 on Monday, down 2.67%. The company reports second-quarter fiscal 2027 earnings on 26 August.

Crypto Briefing
Aug 10th, 2026
Lambda secures $1B credit facility to finance Nvidia chip deployment in AI infrastructure expansion

Lambda, an AI cloud infrastructure company, has closed a $1 billion syndicated senior secured credit facility arranged by J.P. Morgan. This represents a nearly fourfold increase from the $275 million facility Lambda established in August 2025. The capital will fund deployment of next-generation Nvidia AI accelerator servers and data centre expansion. Lambda previously secured a $500 million loan in 2024 collateralised by Nvidia chips. The company raised $480 million in a Series D funding round in February 2025. Lambda has also established a multibillion-dollar commercial partnership with Microsoft focused on building AI infrastructure. Founded in 2012, Lambda now serves tens of thousands of customers with GPU clusters built for AI training and inference workloads.

Newsworthy AI
Aug 4th, 2026
Lambda completes $6.5M secondary share purchase to expand AI cloud infrastructure

Lambda, an AI cloud infrastructure provider, has completed a secondary share purchase of approximately $6.5 million. Aegis Capital Corp. served as the placement agent for the transaction, which involved the sale of existing shares. The transaction provides liquidity for early investors and employees whilst potentially attracting new institutional investors. Founded in 2012 by machine learning engineers, Lambda operates as "The Superintelligence Cloud" and serves tens of thousands of customers, including AI researchers, enterprises, and hyperscalers. Lambda specialises in delivering high-performance supercomputers for AI training and inference. The company's mission is to make compute as accessible as electricity and provide superintelligence capabilities to everyone. The move comes as AI adoption accelerates across industries and demand for computational power continues to surge.

Lambda
May 7th, 2026
Lambda secures $1B credit facility to expand gigawatt-scale AI infrastructure

Lambda, an AI cloud infrastructure provider, has closed a $1 billion senior secured credit facility, nearly quadrupling its original $275 million facility established in August 2025. J.P. Morgan led the oversubscribed arrangement. The multi-tranche facility will fund expansion of Lambda's next-generation NVIDIA AI accelerator servers and data centre capacity. The company operates gigawatt-scale AI factories serving researchers, enterprises and hyperscalers. Founded in 2012, Lambda describes itself as "The Superintelligence Cloud" and serves tens of thousands of customers. CFO Charles Fisher said the financing addresses "unprecedented demand" from AI customers whilst lowering the company's blended cost of capital. Davis Polk & Wardwell represented Lambda, whilst Willkie Farr & Gallagher advised the lenders on the transaction.

Associated Press
May 5th, 2026
Lambda appoints new CEO and chairman to scale AI infrastructure to 3GW by 2030 after $1.5B Series E

Lambda, an AI cloud infrastructure provider, has appointed Michel Combes as CEO and John Donovan as Chairman of the Board, effective May 2026. Co-founder Stephen Balaban will become CTO, whilst co-founder Michael Balaban takes on the Chief Product Officer role. The leadership restructuring follows Lambda's $1.5 billion Series E funding in November 2025. The company has also recently appointed Jerry Hunter as Vice Chairman of Compute Delivery, Charles Fisher as CFO, and David Connolly as CLO. Combes previously served as CEO of Brightspeed, Softbank International, Sprint and Alcatel-Lucent. Donovan held senior positions at AT&T, including CEO of AT&T Communications. The expanded leadership team aims to help Lambda reach 3GW of AI compute under management by 2030, serving frontier labs, enterprises and hyperscalers.