Full-Time

Senior Systems Software Engineer

AI Stack and Performance, DGX Station

Updated on 8/22/2026

Deadline 8/26/26
NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$224k - $356.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Remote in USA + 1 more

More locations: Santa Clara, CA, USA

Remote

Bachelor's, Master's

Category
Software Engineering (1)
Required Skills
Graphics Processing Unit (GPU)
Python
TensorFlow
Neural Networks
CUDA
PyTorch
OpenAI
C/C++

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • A Bachelor’s or Master’s degree in Computer Science or Electrical Engineering, or equivalent experience.
  • At least 12 years of systems software engineering experience with hands-on experience in AI/ML workload optimization, GPU performance analysis, or deep learning infrastructure.
  • Strong proficiency with deep learning frameworks such as PyTorch, TensorFlow, or JAX, including graph execution, operator dispatch, memory management, and custom kernel integration.
  • Experience profiling and optimizing GPU workloads using Nsight Systems, Nsight Compute, CUPTI, or equivalent, with the ability to read GPU traces and translate observations into actionable optimizations.
  • Strong understanding of GPU architecture, including compute units, memory hierarchy, NVLink, multi-GPU scaling, and their impact on AI workload performance.
  • Experience with inference optimization, including quantization with INT8 or FP8, model compilation with TensorRT or torch.compile, batching strategies, and serving frameworks.
  • Proficiency in C, C++, CUDA, and Python, with the ability to read and modify GPU kernels.
Responsibilities
  • Own production readiness of AI applications on DGX Station, including NemoClaw, Hermes agents, NIM microservices, and key customer workloads.
  • Define ready-to-ship criteria, run validation, and close gaps between basic execution and effective performance across single-GPU and multi-GPU configurations.
  • Profile and optimize large language model and deep learning workloads using PyTorch, TensorFlow, and JAX across training and inference on the GB300 Blackwell multi-GPU architecture.
  • Characterize performance across model sizes, batch sizes, precision modes, and GPU scaling to establish benchmarks and identify regressions.
  • Identify bottlenecks in GPU compute, NVLink bandwidth, host memory, PCIe, and CPU–GPU communication.
  • Implement or drive optimizations across kernel tuning, memory placement, NVLink utilization, data pipeline efficiency, and scheduling to increase throughput on DGX Station’s multi-GPU topology.
  • Work with framework, compiler, and GPU architecture teams to improve kernel fusion, graph execution, operator scheduling, and memory management for Blackwell GPUs.
  • Translate DGX Station’s platform-specific constraints and multi-GPU topology into actionable optimization requests for upstream teams.
  • Validate multi-user and concurrent workload scenarios, including simultaneous training jobs, inference serving alongside development, and resource isolation through MIG or time-slicing.
  • Validate the full NVIDIA AI software stack on DGX Station, including the CUDA toolkit, cuDNN, TensorRT, NCCL, Triton Inference Server, DCGM, and DOCA/OFED.
  • Ensure version compatibility, functional correctness, and performance parity with reference data-center configurations.
  • Build and maintain performance benchmarking infrastructure for DGX Station.
  • Track automated regressions across key models, framework versions, and driver updates, and make performance data visible and actionable for general-availability release decisions.
  • Work with product management and OEM/OSV partners to understand target use cases and ensure DGX Station delivers compelling performance.
  • Support customer deployment readiness and field critical issues.
Desired Qualifications
  • Experience optimizing large language model training or inference on multi-GPU NVIDIA systems, including DGX, HGX, or multi-GPU workstations.
  • Contributions to open-source AI frameworks, CUDA libraries, or inference engines.
  • Experience with multi-GPU communication optimization, including NCCL tuning, NVLink utilization, collective operations, and parallel training strategies.
  • A track record of collaborating with compiler and hardware architecture teams to drive kernel fusion, graph optimization, or hardware-specific performance improvements.
  • Experience shipping AI-powered products where application performance on specific hardware was a hard shipping requirement.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Reuters reported only very few H200 chips shipped to China, preserving optionality.
  • NVIDIA's August 2026 financing platforms mobilize outside capital for customers buying its hardware.
  • Poolside and Cloverleaf investments deepen software and power access for future datacenter demand.

What critics are saying

  • U.S. export controls blocked Blackwell shipments to China, shrinking a major demand pool.
  • France's competition authority nears ending its antitrust probe, risking fines and remedies.
  • Circular-financing backlash around $500 billion platforms and Poolside licensing could trigger scrutiny by 2027.

What makes NVIDIA unique

  • NVIDIA controls the AI stack, from Blackwell chips to CUDA software and networking.
  • August 10, 2026 financing partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR broaden distribution.
  • Poolside's Model Factory deal embeds NVIDIA inside model-building workflows, not just accelerators.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-3%
Yahoo Finance
Aug 23rd, 2026
AI infrastructure boom drives $1.7T bond issuance, pushing 30-year Treasury to 19-year high

US corporate bond issuance reached nearly $1.7 trillion year-to-date, up 27% from last year, driven largely by AI infrastructure spending. The surge pushed the 30-year Treasury yield to 5.323% on 18 August, its highest level in 19 years. BMO Capital Markets notes that heavy corporate issuance adds duration supply to fixed-income markets, competing with Treasury bonds for investor capital and lifting yields across the curve. Companies like Vertiv issued $2.1 billion in notes to fund AI buildouts. The dynamic creates a feedback loop: AI-related bond issuance raises long-term interest rates, which in turn increases the discount rate on future earnings that underpin AI company valuations. NVIDIA, with a market capitalisation above $5.26 trillion, sits at the centre of this wave, holding $119 billion in supply commitments and $30 billion in cloud obligations.

Yahoo Finance
Aug 23rd, 2026
Cramer: Nvidia doesn't need anyone but Elon Musk as biggest chip buyer

CNBC host Jim Cramer expressed continued confidence in Nvidia despite the stock's modest 13.7% year-to-date gain in 2026. He suggested the chipmaker doesn't need to invest in AI companies anymore, believing Elon Musk will become its biggest customer. Cramer predicted Musk could purchase Nvidia's entire Vera Rubin production. However, concerns persist about AI spending sustainability, with UBS projecting hyperscaler capital expenditure growth could slow to 25% in 2027 and 6% in 2028. Nvidia's Q1 fiscal 2027 results showed data centre revenue jumping 92% annually to $75 billion, with 75% gross margins. The company guided Q2 revenue to $91 billion, exceeding analyst estimates of $86.84 billion. Bulls project earnings per share could exceed $15 in 2027 and $20 in 2028.

Yahoo Finance
Aug 22nd, 2026
Nvidia partners with Blackstone and 5 firms to mobilise $500B for AI infrastructure

NVIDIA announced partnerships with Blackstone, Apollo, BlackRock, Brookfield, Goldman Sachs and KKR in August 2026 to create AI compute financing platforms targeting over $500 billion in third-party capital for AI infrastructure. Final agreements remain pending. The same month, Blackstone was reportedly evaluating a potential $1.50 billion to $2.00 billion acquisition of Indian renewables platform Blupine Energy from Actis. The moves highlight Blackstone's focus on digital infrastructure and energy transition assets. Analysts note the NVIDIA partnership aligns with Blackstone's existing data centre and private credit commitments, potentially deepening its AI infrastructure financing role. However, concerns about interest rates, deal flow, and market volatility persist. Blackstone's narrative projects $22.5 billion revenue and $9.8 billion earnings by 2029.

Yahoo Finance
Aug 22nd, 2026
Whale Rock dumps 64% of Nvidia stake, shifts $1.5M into AMD shares

Whale Rock Capital Management slashed its Nvidia stake by approximately 64% in the second quarter, reducing its position from 1.04 million shares to 377,204 shares, according to the hedge fund's latest 13F filing. The move appears to be a portfolio rebalancing rather than a retreat from AI semiconductors. During the same period, Whale Rock dramatically increased its Advanced Micro Devices holdings from 69,211 shares to roughly 1.51 million shares, whilst also reducing its Broadcom stake. Nvidia, which designs GPUs and accelerated computing platforms for AI infrastructure, currently trades near $215 per share with a market capitalisation of $5.24 trillion. The company recently reported quarterly revenue growth of 85% year-over-year. The stock trades at a forward price-to-earnings multiple of 25.2 times.

Yahoo Finance
Aug 21st, 2026
Nvidia dominates AI chip market with $75B data center revenue, dwarfing AMD and Qualcomm

Nvidia continues to dominate the AI chip market despite competition from Advanced Micro Devices and Qualcomm, according to recent data. The company's data centre revenue reached $75.2 billion in the first quarter of fiscal 2027, growing 92% year over year. By comparison, AMD's data centre revenue totalled $6.7 billion in its most recent quarter, whilst Qualcomm is targeting $15 billion in data centre revenue by fiscal 2029. Nvidia controls an estimated 74% of the AI inference chip market and holds 80% to 90% of the overall AI chip market, according to Silicon Analysts. With the AI chip market expected to reach $2 trillion by 2030, Nvidia's dominant position suggests significant long-term growth potential.