Full-Time

Senior Software Engineer

TensorRT Edge-LLM

Deadline 8/9/26
NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$152k - $287.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

California, USA + 2 more

More locations: Austin, TX, USA | Santa Clara, CA, USA

Hybrid

Bachelor's, Master's, PhD

Category
AI & Machine Learning (1)
Software Engineering (1)
Required Skills
LLM
CUDA
C/C++
Robotics

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • A Bachelor of Science, Master of Science, or Doctor of Philosophy degree, or equivalent experience, in Computer Science, Electrical Engineering, Computer Engineering, or a closely related field.
  • At least 4 years of relevant software development experience.
  • Deep understanding of transformer models and inference optimization techniques, including quantization, tensor parallelism, or memory-efficient scheduling.
  • Proficiency in modern C++ programming, including C++11, C++14, C++17, and later standards.
  • Familiarity with large language model frameworks and libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer.
  • A track record of strong software design, execution, and cross-functional collaboration.
Responsibilities
  • Develop and evolve an inference framework in modern C++ that extends TensorRT with autoregressive model serving capabilities, including speculative decoding, LoRA, mixture-of-experts, and key-value cache management.
  • Design and implement compiler and runtime optimizations for transformer-based models on constrained, real-time platforms.
  • Collaborate with CUDA, kernel library, compiler, and robotics teams to deliver high-performance, production-ready solutions.
  • Contribute to CUDA kernel and operator development for transformer components such as attention, general matrix multiplication, and mixture-of-experts.
  • Benchmark, profile, and optimize inference performance across embedded and automotive environments.
  • Track the large language model, vision-language model, and vision-language-action ecosystem and incorporate emerging techniques into production software.
Desired Qualifications
  • Development experience or open-source contributions to large language model inference frameworks and libraries such as SGLang, vLLM, or FlashInfer.
  • Proficiency with CUDA, including efficient kernel development, performance profiling, and GPU architecture fundamentals.
  • Prior work on autoregressive large language model serving systems, including speculative decoding or key-value cache management.
  • Familiarity with compiler infrastructure for large language model inference.
  • Exposure to robotics or embedded artificial intelligence pipelines, including optimization for low-latency, resource-constrained systems.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • NVIDIA’s August 10, 2026 financing platforms target over $500 billion for AI infrastructure.
  • Vera Rubin production shipments start this fall, creating a fresh upgrade cycle by 2027.
  • SB Energy’s Ohio campus guarantees exclusive NVIDIA AI compute, securing power-constrained demand.

What critics are saying

  • U.S. export controls keep NVIDIA shut out of China’s largest advanced-AI chip market.
  • Commerce tightened overseas-subsidiary rules on May 31, 2026, crushing workaround sales channels.
  • AMD MI350 and future MI450 racks directly attack NVIDIA pricing and platform leadership in 2026.

What makes NVIDIA unique

  • NVIDIA shipped Vera Rubin into full production by May 31, 2026, ahead of rivals.
  • Blackwell and Rubin plus CUDA lock developers into NVIDIA’s hardware-software stack.
  • August 10, 2026 partnerships with Apollo, BlackRock, and KKR turn compute into finance.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-3%
Yahoo Finance
Aug 23rd, 2026
Cramer: Nvidia doesn't need anyone but Elon Musk as biggest chip buyer

CNBC host Jim Cramer expressed continued confidence in Nvidia despite the stock's modest 13.7% year-to-date gain in 2026. He suggested the chipmaker doesn't need to invest in AI companies anymore, believing Elon Musk will become its biggest customer. Cramer predicted Musk could purchase Nvidia's entire Vera Rubin production. However, concerns persist about AI spending sustainability, with UBS projecting hyperscaler capital expenditure growth could slow to 25% in 2027 and 6% in 2028. Nvidia's Q1 fiscal 2027 results showed data centre revenue jumping 92% annually to $75 billion, with 75% gross margins. The company guided Q2 revenue to $91 billion, exceeding analyst estimates of $86.84 billion. Bulls project earnings per share could exceed $15 in 2027 and $20 in 2028.

Yahoo Finance
Aug 22nd, 2026
Nvidia partners with Blackstone and 5 firms to mobilise $500B for AI infrastructure

NVIDIA announced partnerships with Blackstone, Apollo, BlackRock, Brookfield, Goldman Sachs and KKR in August 2026 to create AI compute financing platforms targeting over $500 billion in third-party capital for AI infrastructure. Final agreements remain pending. The same month, Blackstone was reportedly evaluating a potential $1.50 billion to $2.00 billion acquisition of Indian renewables platform Blupine Energy from Actis. The moves highlight Blackstone's focus on digital infrastructure and energy transition assets. Analysts note the NVIDIA partnership aligns with Blackstone's existing data centre and private credit commitments, potentially deepening its AI infrastructure financing role. However, concerns about interest rates, deal flow, and market volatility persist. Blackstone's narrative projects $22.5 billion revenue and $9.8 billion earnings by 2029.

Yahoo Finance
Aug 22nd, 2026
Whale Rock dumps 64% of Nvidia stake, shifts $1.5M into AMD shares

Whale Rock Capital Management slashed its Nvidia stake by approximately 64% in the second quarter, reducing its position from 1.04 million shares to 377,204 shares, according to the hedge fund's latest 13F filing. The move appears to be a portfolio rebalancing rather than a retreat from AI semiconductors. During the same period, Whale Rock dramatically increased its Advanced Micro Devices holdings from 69,211 shares to roughly 1.51 million shares, whilst also reducing its Broadcom stake. Nvidia, which designs GPUs and accelerated computing platforms for AI infrastructure, currently trades near $215 per share with a market capitalisation of $5.24 trillion. The company recently reported quarterly revenue growth of 85% year-over-year. The stock trades at a forward price-to-earnings multiple of 25.2 times.

Yahoo Finance
Aug 21st, 2026
Nvidia dominates AI chip market with $75B data center revenue, dwarfing AMD and Qualcomm

Nvidia continues to dominate the AI chip market despite competition from Advanced Micro Devices and Qualcomm, according to recent data. The company's data centre revenue reached $75.2 billion in the first quarter of fiscal 2027, growing 92% year over year. By comparison, AMD's data centre revenue totalled $6.7 billion in its most recent quarter, whilst Qualcomm is targeting $15 billion in data centre revenue by fiscal 2029. Nvidia controls an estimated 74% of the AI inference chip market and holds 80% to 90% of the overall AI chip market, according to Silicon Analysts. With the AI chip market expected to reach $2 trillion by 2030, Nvidia's dominant position suggests significant long-term growth potential.

Toscale
Aug 21st, 2026
Nvidia pays $7B for Poolside's Model Factory in licensing deal that avoids acquisition scrutiny

Nvidia has paid $7 billion for a non-exclusive licence to Poolside AI's Model Factory system and taken a minority stake in the company, according to reports from Newcomer, Bloomberg, and The Information. The deal includes a $6 billion licensing fee and a $1 billion investment at a $12 billion pre-money valuation. The transaction follows similar arrangements with Groq ($20 billion) and Enfabrica ($900 million). By structuring deals as licensing agreements rather than acquisitions, Nvidia avoids antitrust scrutiny whilst securing access to critical AI infrastructure and talent. In this case, 109 Poolside employees are transferring to Nvidia, though co-founders Jason Warner and Eiso Kant remain to lead the independent entity. Poolside's existing investors, including Bain Capital Ventures, eBay, and Citi Ventures, are expected to receive the $6 billion licensing payment by end-2027.