Full-Time

Senior System Software Engineer

Agentic Inference, Dynamo

Updated on 9/4/2026

NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$224k - $431.3k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Santa Clara, CA, USA

Remote

Remote within the United States.

Master's, PhD

Category
Software Engineering (1)
Required Skills
LLM
Rust
Python

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • Masters or PhD or equivalent experience
  • 10+ years in Computer Science, Computer Engineering, or related field
  • Ability to work in a fast-paced, agile team environment
  • Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design.
  • Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs
Responsibilities
  • In this role, you will develop open source software to serve inference of trained AI models running on GPUs.
  • Contribute to the development of disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and expand these capabilities to support agentic inference workloads, including long-horizon reasoning, tool calling, and stateful, multi-turn execution.
  • Innovate in inference-state management for long-running agents, including KV- and prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL, to reduce repeated prompt processing, improve latency and token throughput, maximize GPU utilization, and lower per-token and per-task costs for self-hosted LLMs.
  • Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, delivering day-0 support for new models, model-specific request parameters, upstream API compatibility, and stateful Responses API semantics.
  • Balance a variety of objectives: build robust, scalable, high performance software components to support our distributed inference workloads; work with team leads to prioritize features and capabilities; load-balance asynchronous requests across available resources; optimize throughput under latency constraints; and integrate the latest open source technology.
Desired Qualifications
  • Prior contributions to open-source AI inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang).
  • Experience optimizing GPU memory, KV and prefix caches, or high-performance networking for long-context, reasoning, and tool-calling workloads.
  • Understanding of LLM-specific inference challenges for agentic workloads, including context and reasoning-token growth, bursty tool-call-driven traffic, multi-turn state reuse, and scheduling across concurrent trajectories.
  • Prior experience integrating self-hosted LLM serving stacks with agent harnesses such as OpenCode, Codex, Claude Code, and Pi, including compatibility for APIs, streaming, structured outputs, tool calls, and session semantics.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Amazon ordered 2 million GPUs for 2027-2028, validating hyperscaler demand.
  • August 2026 Lancium investment secured 4 GW leased capacity and 15 GW pipeline.
  • Blackwell Ultra and Vera Rubin shipments began; production ships in late 2026.

What critics are saying

  • Taiwan prosecutors indicted NVIDIA staff August 24, 2026 over illegal China server exports.
  • China’s antitrust probe still threatens fines, remedies, and slower mainland sales.
  • China export controls and ASIC substitution can cut NVIDIA off from a trillion-dollar market.

What makes NVIDIA unique

  • CUDA and NVLink lock developers into NVIDIA’s full-stack AI platform.
  • FY2026 revenue hit $215.9 billion, with data center revenue $193.7 billion.
  • Rubin launched February 2026, targeting 10x lower inference costs than Blackwell.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-1%

2 year growth

-2%
Yahoo Finance
Sep 9th, 2026
Nvidia CEO counters Michael Burry's AI chip depreciation bear case with market evidence

Nvidia CEO Jensen Huang has countered hedge fund manager Michael Burry's concerns about the useful life of AI semiconductors. Burry, known for predicting the 2008 subprime market crash, argues that companies overestimate GPU lifespans at six years rather than two to three years, thereby inflating profits. However, Nvidia's Ampere A100 chips, launched in mid-2020, remain in demand six years later. Neocloud CoreWeave recently contracted to rent these processors through 2029, nine years after their introduction. Rental prices for the newer Hopper H100 chips have increased 22% month over month, suggesting continued strong demand. Huang shared this pricing data on social media this week, reinforcing his position against Burry's bearish outlook on AI semiconductor longevity.

Yahoo Finance
Sep 9th, 2026
Nvidia raises Q3 revenue guidance to $108B, but margin squeeze from memory costs dampens investor reaction

Nvidia guided fiscal Q3 2027 revenue to $108 billion, an 18.7% increase from the prior quarter's guidance. The company expects revenue to grow approximately 70% in fiscal 2028, though supply constraints remain a factor. The stock gained 7% following the announcement, adding roughly $14.93 per share. However, rising memory costs driven by AI infrastructure demand are squeezing margins. Nvidia guided gross margin 0.9 percentage points lower for Q3, expecting it to bottom at 71-72% in Q4 2027 before stabilising at 72-73% in fiscal 2028. Operating expenses were also guided higher. The company's data centre business shows strong demand, with its ACIE segment growing 138% year-over-year to $40 billion in Q2 2027.

Yahoo Finance
Sep 9th, 2026
NVIDIA and CU Healthcare Innovation Fund invest in Verily Health's AI platform expansion

Verily Health has secured additional investment from NVIDIA and existing backer CU Healthcare Innovation Fund II. The funding extends Verily's March fundraising round. Current investors include Alphabet, Series X Capital, and UCHealth. The investment supports scaling of Verily's Pre Platform, a data and AI platform helping healthcare organisations deploy AI across research and care. It also backs Verily Me, the company's patient engagement offering. Verily and NVIDIA previously announced a collaboration in October 2025. Researchers using the Pre Platform can leverage NVIDIA AI libraries and B200 GPUs. Verily developed Forecast 1.0, a foundation model integrating genomics and electronic health record data to predict long-term health risks, using NVIDIA's AI stack.

Yahoo Finance
Sep 9th, 2026
NVIDIA stock soars 14,700% in decade as CEO confirms demand outstrips supply through 2028

NVIDIA has delivered a 14,460% return over the past decade, turning $10,000 into approximately $1.45 million. The company's stock currently trades near $226.18, with analysts setting a 12-month price target of $308.85, implying 36.55% upside. NVIDIA reported fiscal Q2 FY27 revenue of $96.22 billion, up 105.85% year-over-year, with data centre revenue hitting $89.02 billion. The company projects Q3 revenue of $108 billion at 74% gross margins. CEO Jensen Huang stated that demand significantly exceeds NVIDIA's current supply capacity. Management guided fiscal 2028 revenue growth of approximately 70%, describing it as supply-constrained. Top-five hyperscaler capital expenditure is projected at nearly $800 billion in 2026 and $1.3 trillion in 2027.

Fortune
Sep 9th, 2026
Nvidia CEO declares AGI achieved with OpenAI's Astra, but markets remain unmoved

Nvidia CEO Jensen Huang declared over the weekend that artificial general intelligence has arrived, citing OpenAI's newest model Astra. Markets responded tepidly: Nvidia shares fell 2% whilst companies leveraged to OpenAI, like CoreWeave and SoftBank, saw gains. Gil Luria of D.A. Davidson suggests Nvidia is "too big to grow" despite reporting $96 billion in quarterly revenue. Economist Basil Halperin argues real interest rates, not stock prices, would signal true AGI. Real rates have risen three to four percentage points since 2021, partly due to AI infrastructure spending, but remain far from levels that would indicate transformative AI impact. Most experts expect AI to add only about half a percentage point to GDP growth. Halperin forecasts the next five years will resemble "the dot-com boom, but twice as fast and twice as hard.