NVIDIA

NVIDIA

Designs GPUs and AI HPC platforms

Software Engineer New Grad - DGX Cloud AI Infrastructure

Full-TimePosted on 9/29/2026Deadline 10/3/26
$108k - $195.5k/yr

+ Equity

Entry
Bachelor's, Master's
Washington, USA+4 more

More locations: Oregon, USA | Austin, TX, USA | Redmond, WA, USA | Santa Clara, CA, USA

Hybrid
Company Historically Provides H1B Sponsorship

About the job

Requirements
  • Bachelor’s or Master’s in Computer Science or a related technical field, or equivalent experience.
  • Experience developing software for AI, high-performance computing, or systems-level applications.
  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
  • Background debugging and scaling distributed systems.
  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Strong Python and C/C++ programming skills.
  • Excellent analytical, debugging, and communication skills, and a collaborative approach across teams.
Responsibilities
  • Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
  • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.
  • Perform root-cause analysis of failures in large distributed environments.
  • Contribute to resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster.
  • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.
  • Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.
  • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.
Desired Qualifications
  • Hands-on experience with NCCL and CUDA-aware distributed execution.
  • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging.
  • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.
  • Experience diagnosing performance jitter.
  • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.

About the company

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • September 28, 2026 board approval added $150 billion buyback authorization, signaling cash generation strength.
  • ByteDance and Alibaba chip-purchase reports on September 27, 2026 reopen a blocked China revenue stream.
  • Over 100 partners joined Open Agent Safety Platform at launch, including Microsoft, JPMorganChase, and SpaceXAI.

What critics are saying

  • China’s September 27, 2026 RTX PRO 5500 reports show export controls still throttle Nvidia sales.
  • OpenAI, Meta, and Google incidents prove agent safety failures can taint Nvidia’s platform credibility fast.
  • A 2027 AI spending downturn would crush data-center demand and expose Nvidia’s customer concentration.

What makes NVIDIA unique

  • CUDA and Blackwell lock developers into Nvidia’s software-hardware stack, unlike commodity chip rivals.
  • September 28, 2026 Open Agent Safety Platform extends Nvidia into agent governance and robotics.
  • BlueField-4 and Vera CPUs let Nvidia control data-center security, networking, and compute together.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

↑ 0%

1 year growth

↓ -1%

2 year growth

↓ -2%
Yahoo Finance
Sep 28th, 2026
Gecko Robotics partners with Nvidia to add safety guardrails for AI in critical infrastructure

Gecko Robotics has partnered with Nvidia to develop safety protocols for AI-powered robots operating in critical infrastructure and military applications. The collaboration focuses on the Open Agent Safety Platform, which aims to keep humans in control whilst enabling AI agents to gather data and direct robotic actions. Gecko Robotics CEO Jake Loosararian emphasised the importance of establishing proper safeguards as AI systems increasingly interact with physical infrastructure. The company describes itself as one of the largest deployers of robotics in real-world settings, with operations affecting critical infrastructure including energy facilities and military assets. The partnership seeks to enable safe deployment of agentic AI tools whilst maintaining human oversight, particularly in high-stakes environments where AI errors could cause significant physical damage or security risks.

Ars Technica
Sep 28th, 2026
Nvidia CEO Jensen Huang emerges as Trump's 'most influential adviser' as China mulls lifting chip export curbs

China is reportedly considering allowing ByteDance and Alibaba to purchase Nvidia's banned RTX Pro 5500 gaming chips, which could be used to power AI models. ByteDance plans to order 1 million chips if approved, securing two quarters of projected sales for Nvidia. Nvidia CEO Jensen Huang has become increasingly influential with President Trump on AI policy. Treasury Secretary Scott Bessent confirmed Trump is "completely aligned with Jensen Huang." Critics worry Nvidia's profit motivations may be steering Trump's thinking on AI regulation. Trump initially tightened export controls but reversed course after meeting Huang at Mar-a-Lago, later lifting restrictions on H200 chips. Neither the US nor China discussed slowing AI development at their recent summit, despite concerns from lawmakers and national security experts about potential risks.

Yahoo Finance
Sep 28th, 2026
Nvidia beats Intel with 55.6% net margin and $96.7B free cash flow in 2026 semiconductor race

Intel reported revenue of $52.9 billion for fiscal year 2025, a slight decline of 0.5% year-on-year, whilst posting a net loss of $267 million. The company is shifting strategy to manufacture chips for external designers through its new foundry model, though shareholder litigation over a government equity deal remains a concern. Nvidia reported revenue of $215.9 billion for the fiscal year ended January 2026, up 65.5% year-on-year, with net income of $120.1 billion. The company dominates the market for graphics processing units used in AI infrastructure but faces risks from customer concentration and export controls. Intel's debt-to-equity ratio stood at 0.4x with negative free cash flow of $4.9 billion. Nvidia maintained a debt-to-equity ratio of 0.1x with free cash flow of $96.7 billion.

Yahoo Finance
Sep 28th, 2026
Nvidia shares rise 3% on launch of AI agent safety platform with OpenShell and Sentry

Nvidia launched a platform designed to keep autonomous AI agents within set operational boundaries, adding a security layer to its infrastructure business. Shares rose 3% on Monday morning. The platform offers two control mechanisms. OpenShell restricts which files, networks, and tools an agent can access. Sentry operates independently on BlueField-4 hardware, monitoring for rule violations and isolating agents within milliseconds if needed. The company is positioning the combination across software, data centres, and robotics applications. The investor thesis centres on free security software potentially drawing more customers into Nvidia's ecosystem, ultimately increasing demand for its processors and networking hardware. Nvidia has not disclosed separate revenue targets for the platform. The company's GF Score stands at 96 out of 100, reflecting strong profitability and growth metrics.

Yahoo Finance
Sep 28th, 2026
Jensen Huang calls OpenAI 'one of the most consequential companies in history

Nvidia CEO Jensen Huang called OpenAI "one of the most consequential companies in history" in a CNBC interview on Monday, praising its leadership and research capabilities. He described working with the AI startup as "a real privilege". Huang's comments followed OpenAI's absence from the roughly 100 companies supporting Nvidia's newly launched OpenShell, an open-source safety platform for AI agents. However, Huang said OpenAI could "use whatever part" of OpenShell it wants. The Nvidia chief described himself as "a responsible optimist" and argued that AI development and safety are complementary rather than competing priorities. He warned against regulatory approaches that could "suffocate the industry" whilst AI continues developing. Nvidia's board authorised a $150 billion increase to its share repurchase programme, bringing total remaining authorisation to $235 billion.