Full-Time

Senior Systems Software Engineer

AI Stack and Performance, DGX Station

Posted on 7/4/2026

Deadline 8/31/26
NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$224k - $356.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Remote in USA + 1 more

More locations: Santa Clara, CA, USA

Hybrid

Bachelor's, Master's

Category
Software Engineering
Required Skills
Graphics Processing Unit (GPU)
Python
TensorFlow
Neural Networks
CUDA
PyTorch
C/C++

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • A BS or MS, or equivalent experience, in Computer Science, Electrical Engineering, or a related field.
  • At least 12 years of systems software engineering experience with hands-on experience in AI/ML workload optimization, GPU performance analysis, or deep learning infrastructure.
  • Strong proficiency with PyTorch, TensorFlow, or JAX, including graph execution, operator dispatch, memory management, and custom kernel integration.
  • Experience profiling and optimizing GPU workloads using Nsight Systems, Nsight Compute, CUPTI, or equivalent, with the ability to read GPU traces and translate observations into actionable optimizations.
  • Strong understanding of GPU architecture, including compute units, memory hierarchy, NVLink, multi-GPU scaling, and their effects on AI workload performance.
  • Experience with inference optimization, including quantization with INT8 or FP8, model compilation with TensorRT or torch.compile, batching strategies, and serving frameworks.
  • Proficiency in C, C++, CUDA, and Python, including the ability to read and modify GPU kernels.
Responsibilities
  • Own production readiness of AI applications on DGX Station, including NemoClaw, Hermes agents, NIM microservices, and key customer workloads.
  • Define ready-to-ship criteria, run validation, and close gaps between an application running and running well across single-GPU and multi-GPU configurations.
  • Profile and optimize LLM and deep learning workloads using PyTorch, TensorFlow, and JAX across training and inference on the GB300 Blackwell multi-GPU architecture.
  • Characterize performance across model sizes, batch sizes, precision modes, and single-GPU versus multi-GPU scaling with NVLink to establish benchmarks and identify regressions.
  • Identify bottlenecks in GPU compute, NVLink bandwidth, host memory, PCIe, and CPU-GPU communication.
  • Implement or drive optimizations across kernel tuning, memory placement, NVLink utilization, data pipeline efficiency, and scheduling to increase throughput on DGX Station's multi-GPU topology.
  • Work with framework, compiler, and GPU architecture teams to improve kernel fusion, graph execution, operator scheduling, and memory management for Blackwell GPUs.
  • Translate DGX Station's platform-specific constraints and multi-GPU topology into actionable optimization requests for upstream teams.
  • Validate multi-user and concurrent workload scenarios, including simultaneous training jobs, inference serving alongside development, and resource isolation through MIG or time-slicing.
  • Validate the full NVIDIA AI software stack on DGX Station, including the CUDA toolkit, cuDNN, TensorRT, NCCL, Triton Inference Server, DCGM, and DOCA/OFED.
  • Ensure version compatibility, functional correctness, and performance parity with reference data center configurations.
  • Build and maintain performance benchmarking infrastructure for DGX Station.
  • Track automated regressions across key models, framework versions, and driver updates, and make performance data visible and actionable for GA release decisions.
  • Work with product management and OEM/OSV partners to understand target use cases and ensure DGX Station delivers compelling performance.
  • Support customer deployment readiness and field critical issues.
Desired Qualifications
  • Experience optimizing LLM training or inference on multi-GPU NVIDIA systems, including DGX, HGX, or multi-GPU workstations.
  • Contributions to open-source AI frameworks, CUDA libraries, or inference engines.
  • Experience with multi-GPU communication optimization, including NCCL tuning, NVLink utilization, collective operations, and parallel training strategies.
  • A track record of collaborating with compiler and hardware architecture teams to drive kernel fusion, graph optimization, or hardware-specific performance improvements.
  • Experience shipping AI-powered products where application performance on specific hardware was a hard shipping requirement.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Nvidia reported $96.2 billion Q2 revenue and 106% growth, showing brutal demand.
  • The $500 billion financing push widens access to Nvidia compute across AI builders.
  • DLSS 5 launched in NBA 2K27 on RTX 50, reinforcing consumer ecosystem loyalty.

What critics are saying

  • Reuters said Nvidia paused revenue-sharing deals August 27, exposing shaky financing economics.
  • OpenAI’s Broadcom Jalapeño chip matches Blackwell, threatening Nvidia’s inference monopoly in 2026.
  • U.S. and Taiwan enforcement against China-bound server exports keeps shrinking Nvidia’s addressable market.

What makes NVIDIA unique

  • Blackwell and Rubin still define the AI stack, from chips to networking to software.
  • DLSS 5 shipped September 2026, keeping GeForce visually ahead in consumer gaming.
  • Nvidia’s financing platforms with Apollo, BlackRock, and KKR deepen customer lock-in.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-1%

2 year growth

-2%
Yahoo Finance
Sep 4th, 2026
Oregon couple's Nvidia stake now exceeds home, 401(k)s combined — posing biggest threat to retirement

An Oregon couple in their late 50s holds an Nvidia position now worth more than their home, 401(k)s, and all other assets combined. The stock has surged 864% over five years to around $221 per share. While Nvidia reported $96 billion in quarterly revenue with 75% gross margins, the concentrated position poses severe retirement risk. With a beta of 2.2, the stock has experienced double-digit drops following earnings despite strong results, including an 11% thirty-day decline after Q4 FY26. A 50% to 70% drawdown near retirement could force the couple to sell depressed shares for living expenses. Advisers recommend reducing concentration through staged selling across tax years, redirecting savings to index funds, or donating appreciated shares to donor-advised funds to avoid a single massive tax hit.

Yahoo Finance
Sep 3rd, 2026
Nvidia shifts from AI infrastructure seller to financier as it backs $500B in customer funding

Nvidia has shifted its business model from primarily selling AI infrastructure to financing its customers' purchases. The company has invested nearly $50 billion in AI labs and arranged financing platforms targeting over $500 billion in outside capital. Management now describes Nvidia as a "full-stack AI factory platform" rather than the "data centre-scale AI infrastructure company" label used two years ago. The chip and networking segment still generates roughly $193.5 billion annually, about 64% of total revenue. The financing approach is clearest in Nvidia's NeoCloud arrangement, where it provides take-or-pay commitments on data centre capacity to help secure project funding, then shares rental revenue above a floor price. Days sales outstanding rose to 60 days in fiscal Q2 2027 due to extended payment terms for large customers. Inventory increased to $32 billion ahead of the Vera Rubin launch. Management forecast 70% revenue growth for fiscal 2028, constrained by supply.

Yahoo Finance
Sep 3rd, 2026
Nvidia eyes $312B from $2T cloud backlog as hyperscalers boost spending to $1.3T by 2027

Nvidia reported fiscal Q2 revenue of $96.2 billion, up 106% year-over-year, with data centre revenue rising 117% to $89.0 billion. CFO Colette Kress disclosed that cloud industry backlog now exceeds $2 trillion, with the top five hyperscalers expected to spend nearly $800 billion in 2026 and $1.3 trillion in 2027. Hyperscaler revenue reached $48.7 billion in Q2, more than double the prior year's $24.2 billion. At a $195 billion annual run rate, this represents approximately 24% of the projected $800 billion hyperscaler spending in 2026. The comparison carries caveats: Nvidia's fiscal year ends in late January, and its hyperscaler category includes more than five customers, suggesting the actual share runs somewhat lower.

Yahoo Finance
Sep 3rd, 2026
Nvidia's DLSS 5 debuts in NBA 2K27 with AI-generated graphics on single GPU

Nvidia is rolling out DLSS 5, its AI-powered graphics technology, starting 3 September with NBA 2K27. The company calls it its largest computer-graphics advance since introducing real-time ray tracing in 2018. DLSS 5 uses what Nvidia terms 3D-Guided Neural Rendering. The neural model analyses scene objects to add realistic lighting, materials, shadows, and visual details. The technology's hardware requirements have dropped dramatically. When Nvidia demonstrated DLSS 5 in March, it required two flagship GeForce RTX 5090 graphics cards. Performance has improved fivefold, and the technology now runs on a single RTX 50-series GPU, including the RTX 5060. Nvidia states nearly every RTX 50-series GPU can run NBA 2K27 at 1080p with ray tracing and Ultra settings using DLSS 5.

Yahoo Finance
Sep 3rd, 2026
Cramer urges Nvidia to launch $500B buyback as chipmaker becomes 'banker' for AI buildout

Jim Cramer has called for NVIDIA to authorise a $500 billion share buyback programme, representing roughly 10% of the company, arguing that the stock does not reflect its earnings power. The suggestion comes as NVIDIA increasingly acts as a financier for AI infrastructure buildouts. NVIDIA reported fiscal Q2 2027 revenue of $96.2 billion, up 106% year-on-year, with data centre revenue reaching $89 billion. The company repurchased 203 million shares for $39.8 billion in the first half of fiscal 2027. Management expects approximately 70% revenue growth in fiscal 2028. However, NVIDIA faces mounting risks. Cramer noted that many companies NVIDIA backs are "non-investment grade". Reuters reported NVIDIA paused certain revenue-sharing agreements with AI cloud companies in August. The company also faces competition from Broadcom and custom chips developed by hyperscalers, alongside exposure to China's export restrictions.

INACTIVE