Full-Time

Senior Systems Software Engineer

Kubernetes Node Lifecycle, DGX Cloud

NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$184k - $356.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Seattle, WA, USA + 1 more

More locations: Santa Clara, CA, USA

In Person

On-site roles in Santa Clara, CA or Seattle, WA.

Bachelor's, Master's

Category
DevOps & Infrastructure (1)
Software Engineering (1)
Required Skills
Packer
Kubernetes
Microsoft Azure
Python
AWS
Go
Google Cloud Platform

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • 8 years of experience with a background in systems software, cloud infrastructure, or Kubernetes node engineering
  • Bachelor’s or Master’s degree in Engineering (Electrical, Computer Engineering, Computer Science) or equivalent experience
  • Deep expertise in Cluster API (CAPI), including provider development and full machine lifecycle from provisioning to deletion
  • Extensive experience with OS image build pipelines, node image packaging, and delivery systems for Kubernetes nodes (for example image-builder, containerd, cloud-init, packer)
  • Practical experience with bring-your-own-node models and integrating diverse hardware into live Kubernetes environments, including large-scale nodepool lifecycle management and upgrades
  • Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle
  • Experience with node image security, including vulnerability scanning, patch automation, and compliance gating as part of image build pipelines
  • Proficiency in Golang and/or Python, and hands-on experience with at least one major public cloud provider (GCP, AWS, Azure, OCI or equivalent)
Responsibilities
  • Direct the building and refinement of CAPI providers for NVIDIA Kubernetes Engine, maintaining steady, consistent, and scalable node provisioning across DGX Cloud and NCP environments
  • Develop and maintain bring-your-own-node workflows that allow customers to integrate different NVIDIA hardware into NKE clusters while ensuring high operational consistency
  • Coordinate OS image generation, packaging, deployment, and update processes for NKE nodes. Ensure images are fine-tuned for NVIDIA GPU workloads and satisfy enterprise- and cloud-grade security and compliance criteria
  • Develop and sustain node image hardening pipelines, incorporating CIS benchmarks, automated CVE remediation, and promotion gates connected to security posture
  • Develop and maintain automated test suites for node images. These tests verify accuracy across Kubernetes versions and NVIDIA hardware configurations. This process occurs prior to production deployment and facilitates continuous validation through modern CI/CD pipelines
  • Handle nodepool lifecycle at scale, including provisioning, upgrades, drain and cordon workflows, and seamless node replacement across very large clusters with diverse NVIDIA hardware
  • Examine, resolve, and determine underlying causes of node-layer faults in production NKE clusters, such as those involving image configuration, driver packaging, kubelet operation, and hardware activation, and review and optimize the node layer in real-world high-scale scenarios
  • Partner with upstream communities including Cluster API, Kubernetes, and CNCF projects to establish node provisioning and lifecycle standards in accordance with NKE requirements. Communicate your progress and findings at internal and external gatherings such as KubeCon and GTC
Desired Qualifications
  • Direct experience building or maintaining node image pipelines for a hyperscaler Kubernetes distribution (GKE, EKS, AKS, OKE, or equivalent)
  • Experience with supply chain security and hardening for node images, including image signing, provenance attestation, SBOM generation, CIS benchmark consistency, and automated CVE remediation
  • Experience with automated node provisioning and optimal sizing at scale (for example Karpenter, GKE NAP or similar) and how these interact with GPU workload scheduling
  • Strong operational experience working with immutable OS image distributions (such as Flatcar, Bottlerocket, Azure Linux) and debugging node-layer failures in large Kubernetes clusters
  • Proven background of upstream contributions to Cluster API, Kubernetes or related CNCF projects, combined with excellent communication and interpersonal abilities

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • NVIDIA’s August 10, 2026 financing platforms target over $500 billion for AI infrastructure.
  • Vera Rubin production shipments start this fall, creating a fresh upgrade cycle by 2027.
  • SB Energy’s Ohio campus guarantees exclusive NVIDIA AI compute, securing power-constrained demand.

What critics are saying

  • U.S. export controls keep NVIDIA shut out of China’s largest advanced-AI chip market.
  • Commerce tightened overseas-subsidiary rules on May 31, 2026, crushing workaround sales channels.
  • AMD MI350 and future MI450 racks directly attack NVIDIA pricing and platform leadership in 2026.

What makes NVIDIA unique

  • NVIDIA shipped Vera Rubin into full production by May 31, 2026, ahead of rivals.
  • Blackwell and Rubin plus CUDA lock developers into NVIDIA’s hardware-software stack.
  • August 10, 2026 partnerships with Apollo, BlackRock, and KKR turn compute into finance.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-3%
Yahoo Finance
Aug 23rd, 2026
Cramer: Nvidia doesn't need anyone but Elon Musk as biggest chip buyer

CNBC host Jim Cramer expressed continued confidence in Nvidia despite the stock's modest 13.7% year-to-date gain in 2026. He suggested the chipmaker doesn't need to invest in AI companies anymore, believing Elon Musk will become its biggest customer. Cramer predicted Musk could purchase Nvidia's entire Vera Rubin production. However, concerns persist about AI spending sustainability, with UBS projecting hyperscaler capital expenditure growth could slow to 25% in 2027 and 6% in 2028. Nvidia's Q1 fiscal 2027 results showed data centre revenue jumping 92% annually to $75 billion, with 75% gross margins. The company guided Q2 revenue to $91 billion, exceeding analyst estimates of $86.84 billion. Bulls project earnings per share could exceed $15 in 2027 and $20 in 2028.

Yahoo Finance
Aug 22nd, 2026
Nvidia partners with Blackstone and 5 firms to mobilise $500B for AI infrastructure

NVIDIA announced partnerships with Blackstone, Apollo, BlackRock, Brookfield, Goldman Sachs and KKR in August 2026 to create AI compute financing platforms targeting over $500 billion in third-party capital for AI infrastructure. Final agreements remain pending. The same month, Blackstone was reportedly evaluating a potential $1.50 billion to $2.00 billion acquisition of Indian renewables platform Blupine Energy from Actis. The moves highlight Blackstone's focus on digital infrastructure and energy transition assets. Analysts note the NVIDIA partnership aligns with Blackstone's existing data centre and private credit commitments, potentially deepening its AI infrastructure financing role. However, concerns about interest rates, deal flow, and market volatility persist. Blackstone's narrative projects $22.5 billion revenue and $9.8 billion earnings by 2029.

Yahoo Finance
Aug 22nd, 2026
Whale Rock dumps 64% of Nvidia stake, shifts $1.5M into AMD shares

Whale Rock Capital Management slashed its Nvidia stake by approximately 64% in the second quarter, reducing its position from 1.04 million shares to 377,204 shares, according to the hedge fund's latest 13F filing. The move appears to be a portfolio rebalancing rather than a retreat from AI semiconductors. During the same period, Whale Rock dramatically increased its Advanced Micro Devices holdings from 69,211 shares to roughly 1.51 million shares, whilst also reducing its Broadcom stake. Nvidia, which designs GPUs and accelerated computing platforms for AI infrastructure, currently trades near $215 per share with a market capitalisation of $5.24 trillion. The company recently reported quarterly revenue growth of 85% year-over-year. The stock trades at a forward price-to-earnings multiple of 25.2 times.

Yahoo Finance
Aug 21st, 2026
Nvidia dominates AI chip market with $75B data center revenue, dwarfing AMD and Qualcomm

Nvidia continues to dominate the AI chip market despite competition from Advanced Micro Devices and Qualcomm, according to recent data. The company's data centre revenue reached $75.2 billion in the first quarter of fiscal 2027, growing 92% year over year. By comparison, AMD's data centre revenue totalled $6.7 billion in its most recent quarter, whilst Qualcomm is targeting $15 billion in data centre revenue by fiscal 2029. Nvidia controls an estimated 74% of the AI inference chip market and holds 80% to 90% of the overall AI chip market, according to Silicon Analysts. With the AI chip market expected to reach $2 trillion by 2030, Nvidia's dominant position suggests significant long-term growth potential.

Toscale
Aug 21st, 2026
Nvidia pays $7B for Poolside's Model Factory in licensing deal that avoids acquisition scrutiny

Nvidia has paid $7 billion for a non-exclusive licence to Poolside AI's Model Factory system and taken a minority stake in the company, according to reports from Newcomer, Bloomberg, and The Information. The deal includes a $6 billion licensing fee and a $1 billion investment at a $12 billion pre-money valuation. The transaction follows similar arrangements with Groq ($20 billion) and Enfabrica ($900 million). By structuring deals as licensing agreements rather than acquisitions, Nvidia avoids antitrust scrutiny whilst securing access to critical AI infrastructure and talent. In this case, 109 Poolside employees are transferring to Nvidia, though co-founders Jason Warner and Eiso Kant remain to lead the independent entity. Poolside's existing investors, including Bain Capital Ventures, eBay, and Citi Ventures, are expected to receive the $6 billion licensing payment by end-2027.