Full-Time

Senior System Architect

Infrastructure Reliability

NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

No salary listed

Company Historically Provides H1B Sponsorship

Austin, TX, USA + 4 more

More locations: Redmond, WA, USA | Santa Clara, CA, USA | Durham, NC, USA | Westford, MA, USA

Hybrid

Hybrid role; some on-site days at listed U.S. locations.

Bachelor's, Master's, PhD

Category
DevOps & Infrastructure (1)
Required Skills
Kubernetes
Python
CUDA
Machine Learning
C/C++

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • Distributed Systems Mastery: BS, MS, or PhD in Computer Science or Electrical Engineering (or equivalent experience) with 6+ years in systems programming.
  • Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments.
  • CPU Architecture Deep-Dive: Expert knowledge of x86/ARM node-level metrics: IPC (Instructions Per Cycle), cache contention, NUMA imbalance, and hardware interrupts.
  • Programming Proficiency: Strong C++ and Python skills, with the ability to build high-performance daemons that monitor system health without impacting workload performance.
  • Scale Experience: Familiarity with cluster resource managers (Slurm, LSF, or Kubernetes) and how they manage job lifecycle and signal propagation.
Responsibilities
  • Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure.
  • Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs.
  • Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters.
  • Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams.
  • Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs.
Desired Qualifications
  • Low-Level Diagnostics: Expert knowledge of the Linux kernel and its error-reporting interfaces (/dev/mcelog, dmesg, journald). Understand how the kernel handles hardware exceptions and memory faults.
  • GPU Infrastructure Proficiency: Deep experience with the NVIDIA DCGM (Data Center GPU Manager) and NVIDIA Management Library (NVML) for monitoring device health and capturing state-dumps.
  • Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns.
  • Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • NVIDIA mobilized over $500 billion for AI infrastructure on August 10, 2026.
  • NetSense pilot deployments start late 2026, creating new edge-AI revenue beyond data centers.
  • Nemotron 3.5 Lightning targets inference, where Gartner says spending reaches $23.3 billion this year.

What critics are saying

  • Six Wall Street partners concentrate NVIDIA demand into a credit-sensitive financing machine.
  • Nemotron routing tools commoditize inference, inviting margin pressure from open models and rivals.
  • If hyperscaler spending stalls in 2027, NVIDIA's order book and valuation reset fast.

What makes NVIDIA unique

  • NVIDIA AI Aerial turned Verizon 5G into drone sensing with Lockheed Martin on August 12, 2026.
  • Nemotron 3.5 Lightning and NeMo Switchyard bundle models, routing, and hardware into one stack.
  • The August 10, 2026 financing platform makes NVIDIA compute an investable infrastructure asset.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-3%
Yahoo Finance
Aug 13th, 2026
Verizon, Nvidia and Lockheed Martin use 5G to track drones in real time

Verizon Communications has partnered with Lockheed Martin, Nvidia, Keysight Technologies, ODC and Astris AI to demonstrate drone-tracking technology using 5G networks. The system, called NetSense, was tested in Miami in July. NetSense combines Verizon's existing 5G spectrum with Nvidia's AI Aerial platform, ODC's AI-native Radio Access Network software, and Keysight's radio-frequency simulation technology. The system analyses radio-frequency disturbances to identify drones, predict flight paths and provide real-time alerts. The demonstration showed NetSense could detect and track drones without requiring changes to existing cellular infrastructure. The technology is designed for airports, power plants, stadiums, schools, hospitals and other critical infrastructure, as well as government monitoring. Separately, Morningstar named Verizon among its top 10 dividend stocks.

Yahoo Finance
Aug 12th, 2026
Intel's market value soars 474% in a year vs Nvidia's 24% — but there's a catch

Intel's market value surged 474% over the past year to roughly $510 billion, whilst Nvidia's grew 24% to near $5.4 trillion. Intel's stock price rose approximately 400% from around $20 to about $101, with additional gains from share count increases of roughly one-sixth. Intel's share count grew from about 4.4 billion to over 5 billion shares, largely through stock issued to the US government under the CHIPS Act agreement. The company's revenue rose 25% year-over-year last quarter, its fastest growth in nearly 15 years, though it posted an $11.3 billion trailing-12-month loss. Nvidia generated $159.6 billion in trailing earnings, more than double the prior year, on $253 billion of revenue. Its price-to-earnings ratio now sits near 34.

Yahoo Finance
Aug 12th, 2026
Nvidia backs open-source AI to boost hardware demand as inference spending hits $23.3B

Nvidia has released its open-source Nemotron 3.5 Lightning model, reinforcing CEO Jensen Huang's public support for open-source artificial intelligence. The move positions the chip maker opposite closed-model proponents like OpenAI and Anthropic in the AI development debate. The strategy serves Nvidia's commercial interests by directing enterprise spending towards hardware rather than expensive proprietary software. "What Nvidia is doing is removing that cost, which means [enterprise clients] have more money for hardware," said Bill Wong, AI research fellow at Info-Tech Research Group. Global AI inference spending is projected to reach $23.3 billion this year, surpassing training expenditure for the first time, according to Gartner. Nvidia holds 90% of the AI training market and has increased its inference market share from 66% to 74% year-on-year.

The Register
Aug 12th, 2026
Nvidia launches router to slash AI costs by 74% using model switching

Nvidia has introduced NeMo Switchyard, a router designed to reduce enterprise AI costs by directing prompts to different models based on cost, latency, or quality requirements. The platform routes requests between expensive proprietary models and cheaper, smaller alternatives, potentially cutting job completion costs by 74% compared to using Claude Opus 4 alone, with approximately six-point accuracy trade-off. Switchyard functions as a proxy between inference servers and models, optimising which model handles each task. Nvidia suggests using smaller, specialised models for simple tasks like generating title cards, whilst reserving larger models for complex work. The company has developed application-specific models, including Nemotron Parse for PDF processing. Similar routing approaches have been adopted by OpenAI and AT&T. The telecommunications giant reportedly saved 80-90% in certain applications by switching to open-weight models, which now power 25% of its AI workloads.

Yahoo Finance
Aug 12th, 2026
Nvidia partners with BlackRock, Goldman Sachs on $500B AI data centre financing deal

Nvidia has partnered with major financial firms including Apollo, BlackRock, Blackstone, and Goldman Sachs to establish a $500 billion facility for AI data centre infrastructure. The initiative will focus on debt financing to provide computing access for Nvidia's largest customers. Nvidia reported quarterly revenue of $81 billion, up 85% year over year, putting its revenue run rate near $330 billion. The company projects $91 billion in revenue for the current quarter. Nvidia maintains financial relationships and ownership stakes in numerous AI companies including Anthropic, OpenAI, Intel, Coreweave, Nokia, Synopsys, and Marvell. The infrastructure partnership is expected to generate additional revenue for Nvidia, supporting its continued top-line growth of over 60% quarterly.