Full-Time

Senior Software Engineer

Distributed Systems Engineer, EDA Infrastructure

Updated on 9/9/2026

Deadline 9/12/26
NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$152k - $287.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Remote in USA + 5 more

More locations: Washington, USA | California, USA | Austin, TX, USA | Durham, NC, USA | Westford, MA, USA

Hybrid

Bachelor's

Category
DevOps & Infrastructure (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Python
Distributed Systems
Data Structures & Algorithms
Computer Networking
Go
Observability
Linux/Unix

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • At least 5 years of software engineering or infrastructure engineering experience supporting large-scale production systems.
  • A Bachelor of Science degree in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
  • Strong programming experience in Go or Python, including data structures, algorithms, testing, and software design.
  • Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
  • Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication skills and the ability to work effectively across teams, organizations, and geographic regions.
  • A systematic approach to problem solving, a strong sense of ownership, and an emphasis on reducing operational toil.
Responsibilities
  • Design and build platforms that automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and remediation systems that improve the reliability, availability, and utilization of EDA compute environments.
  • Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
  • Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
  • Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and return unhealthy systems to service.
  • Work with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical chip-design workloads.
  • Participate in incident response, root-cause analysis, capacity planning, and continuous improvement of production services.
Desired Qualifications
  • Experience designing or operating large-scale EDA or high-performance computing infrastructure.
  • Deep knowledge of Linux, GPU and CPU server architecture, networking, storage, and bare-metal lifecycle management.
  • Hands-on experience with workload schedulers and cluster-management platforms such as Slurm, LSF, Kubernetes, or Bright Cluster Manager.
  • Experience supporting EDA applications, license-management systems, high-throughput batch workloads, or semiconductor design workflows.
  • Experience building automated health checks, break-fix remediation, firmware and operating-system upgrade workflows, or node-provisioning systems.
  • A track record of improving infrastructure reliability, utilization, and recovery time through production-quality automation.
  • Experience operating infrastructure across multiple data centers or heterogeneous hardware environments.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Amazon ordered 2 million GPUs for 2027-2028, validating hyperscaler demand.
  • August 2026 Lancium investment secured 4 GW leased capacity and 15 GW pipeline.
  • Blackwell Ultra and Vera Rubin shipments began; production ships in late 2026.

What critics are saying

  • Taiwan prosecutors indicted NVIDIA staff August 24, 2026 over illegal China server exports.
  • China’s antitrust probe still threatens fines, remedies, and slower mainland sales.
  • China export controls and ASIC substitution can cut NVIDIA off from a trillion-dollar market.

What makes NVIDIA unique

  • CUDA and NVLink lock developers into NVIDIA’s full-stack AI platform.
  • FY2026 revenue hit $215.9 billion, with data center revenue $193.7 billion.
  • Rubin launched February 2026, targeting 10x lower inference costs than Blackwell.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-1%

2 year growth

-2%
Business Insider
Sep 6th, 2026
Nvidia's Huang declares 'AGI has arrived' as OpenAI unveils Astra model

Nvidia CEO Jensen Huang declared that "AGI has arrived" whilst congratulating OpenAI on its newest model, Astra, released last Thursday. Astra was trained on Nvidia's chips and is described by OpenAI as the world's "most intelligent and aligned model". AGI, or artificial general intelligence, refers to AI systems that match or surpass human intelligence. OpenAI defines it as "highly autonomous systems that outperform humans at most economically valuable work". OpenAI president Greg Brockman said people will likely look back and think AGI was created "about this time". However, OpenAI CEO Sam Altman recently called AGI "a very poorly defined term" and "an irrelevant marketing term". Nvidia reported $96.2 billion in quarterly revenue in August, more than double the previous year, driven largely by demand for its AI chips.

Yahoo Finance
Sep 6th, 2026
Nvidia CEO Jensen Huang declares AGI achieved using 100,000 chip boxes as company posts $89B quarterly AI sales

NVIDIA CEO Jensen Huang claimed artificial general intelligence has arrived, citing OpenAI's latest ChatGPT model trained on his company's chips. NVIDIA now sells 72 chips in one box, wired to function as a unified system. Huang stated roughly 100,000 of these boxes trained OpenAI's new model, though he initially posted 300,000 before deleting and revising the figure without explanation. NVIDIA sold $89 billion of AI computers in three months, more than double the previous year. However, OpenAI has not confirmed AGI's arrival, and its president Greg Brockman stopped short of making such claims. A previously announced $100 billion partnership between the firms was never signed, and OpenAI has reportedly been buying fewer NVIDIA chips.

Omega Technology Solutions Group, Inc.
Sep 6th, 2026
NVIDIA invests in Lancium to secure 15+ GW AI data centre pipeline

NVIDIA has taken a strategic investment in Lancium, a Blackstone-backed infrastructure company, to deploy its AI platform across power-ready data centres. The deal provides NVIDIA access to 4 gigawatts of leased capacity and a development pipeline exceeding 15 gigawatts. The partnership addresses a critical bottleneck in AI infrastructure, where compute demand has outpaced available power-ready capacity at gigawatt scale. NVIDIA's DSX MaxLPS technology enables up to 40% more GPU density within the same power budget, improving deployment economics. Lancium is backed by Blackstone Energy Transition Partners and Blackstone Multi-Asset Investing. The investment signals NVIDIA is moving beyond hardware sales to secure physical capacity for customers, recognising that compute demand requires corresponding power and space infrastructure. The arrangement allows NVIDIA customers faster access to large-scale infrastructure built specifically for demanding AI workloads.

Yahoo Finance
Sep 6th, 2026
Nvidia and Broadcom forecast $230B AI chip revenue surge by 2028

AI stocks are predicted to remain strong investments over the next five years, with recent corporate guidance supporting this outlook. Broadcom reported 86% year-over-year revenue growth in its fiscal 2026 third quarter, with AI semiconductor revenue up 221%. The chipmaker projects AI chip revenue will double to $115 billion in fiscal 2027, then double again to $230 billion in fiscal 2028. Nvidia anticipates 70% year-over-year revenue growth in its fiscal 2028, citing supply chain issues as a limiting factor. Hyperscalers including Amazon, Microsoft, and Alphabet are generating substantial returns from their AI infrastructure investments through their cloud platforms. Investors are increasingly looking beyond chipmakers to smaller AI stocks for diversification opportunities.

Yahoo Finance
Sep 6th, 2026
Nvidia's $12.9B Hugging Face acquisition shows IPOs becoming optional for startups

Nvidia confirmed Thursday it will acquire AI model distribution platform Hugging Face for $12.9 billion. The company had reached $150 million in annualised revenue and raised nearly $395 million from investors including Amazon, Intel, Sequoia, and Coatue. The deal exemplifies how IPOs are becoming optional rather than obligatory for venture-backed companies. According to PitchBook research, the public offering is now "a tool for a specific problem" rather than an expected destination. Recent public market struggles support this shift. Chime went public at a steep markdown, whilst Figma has traded below its offer price for most of 2026. Only companies with capital needs too large for private buyers, such as OpenAI and Anthropic, now require public listings. Other successful exits include Stripe's acquisition of OpenRouter and SpaceX buying Cursor.