Full-Time

Deep Learning Communication Architect Senior

NVIDIA

NVIDIA

10,001+ employees

Designs GPUs and AI HPC platforms

Compensation Overview

$184k - $287.5k/yr

+ Equity

Company Historically Provides H1B Sponsorship

Austin, TX, USA + 1 more

More locations: Santa Clara, CA, USA

In Person

Bachelor's, Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
Neural Networks
CUDA
PyTorch
C/C++

Get referred to NVIDIA

See people who can refer or advise you

Requirements
  • A Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or a closely related field or equivalent experience.
  • 6+ years of experience in building DNNs, scaling of DNNs, parallelism of DNN frameworks, or deep learning training and inference workloads.
  • Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
  • Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and Fully Sharded Data Parallelism.
  • Understanding of emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
  • Proficiency in developing code for one or more deep neural network training and inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.
Responsibilities
  • The software architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of thousands of nodes.
  • Optimizing communication performance: Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
  • Designing efficient communication protocols: Develop and implement communication algorithms and protocols tailored for deep learning workloads, minimizing communication overhead and latency.
  • Hardware and software co-craft: Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC, NVSHMEM).
  • Exploring innovative communication technologies: Research and evaluate new communication technologies and techniques to enhance the performance and scalability of deep learning systems.
  • Developing and implementing solutions: Build proofs-of-concept, conduct experiments, and perform quantitative modeling to validate and deploy new communication strategies.
Desired Qualifications
  • Prior contributions to one or more DNN training and Inference frameworks as part of your previous work experience.
  • Deep understanding and contributions to the scaling of LLMs on large-scale systems.

NVIDIA designs and manufactures graphics processing units (GPUs) and computing platforms used for gaming, data centers, and artificial intelligence. These products work by using parallel processing to handle complex mathematical calculations much faster than standard computer processors, supported by a software ecosystem that allows developers to build and run AI models. Unlike competitors that may focus solely on hardware, NVIDIA integrates its chips with specialized software and cloud services to create a complete environment for high-performance tasks. The company’s goal is to provide the underlying technology necessary to power advanced computing, from realistic video game graphics to autonomous vehicles and large-scale data analysis.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1993

Get referred to NVIDIA

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • NVIDIA mobilized over $500 billion for AI infrastructure on August 10, 2026.
  • NetSense pilot deployments start late 2026, creating new edge-AI revenue beyond data centers.
  • Nemotron 3.5 Lightning targets inference, where Gartner says spending reaches $23.3 billion this year.

What critics are saying

  • Six Wall Street partners concentrate NVIDIA demand into a credit-sensitive financing machine.
  • Nemotron routing tools commoditize inference, inviting margin pressure from open models and rivals.
  • If hyperscaler spending stalls in 2027, NVIDIA's order book and valuation reset fast.

What makes NVIDIA unique

  • NVIDIA AI Aerial turned Verizon 5G into drone sensing with Lockheed Martin on August 12, 2026.
  • Nemotron 3.5 Lightning and NeMo Switchyard bundle models, routing, and hardware into one stack.
  • The August 10, 2026 financing platform makes NVIDIA compute an investable infrastructure asset.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Company Equity

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

-2%

2 year growth

-3%
Yahoo Finance
Aug 13th, 2026
Verizon, Nvidia and Lockheed Martin use 5G to track drones in real time

Verizon Communications has partnered with Lockheed Martin, Nvidia, Keysight Technologies, ODC and Astris AI to demonstrate drone-tracking technology using 5G networks. The system, called NetSense, was tested in Miami in July. NetSense combines Verizon's existing 5G spectrum with Nvidia's AI Aerial platform, ODC's AI-native Radio Access Network software, and Keysight's radio-frequency simulation technology. The system analyses radio-frequency disturbances to identify drones, predict flight paths and provide real-time alerts. The demonstration showed NetSense could detect and track drones without requiring changes to existing cellular infrastructure. The technology is designed for airports, power plants, stadiums, schools, hospitals and other critical infrastructure, as well as government monitoring. Separately, Morningstar named Verizon among its top 10 dividend stocks.

Yahoo Finance
Aug 12th, 2026
Intel's market value soars 474% in a year vs Nvidia's 24% — but there's a catch

Intel's market value surged 474% over the past year to roughly $510 billion, whilst Nvidia's grew 24% to near $5.4 trillion. Intel's stock price rose approximately 400% from around $20 to about $101, with additional gains from share count increases of roughly one-sixth. Intel's share count grew from about 4.4 billion to over 5 billion shares, largely through stock issued to the US government under the CHIPS Act agreement. The company's revenue rose 25% year-over-year last quarter, its fastest growth in nearly 15 years, though it posted an $11.3 billion trailing-12-month loss. Nvidia generated $159.6 billion in trailing earnings, more than double the prior year, on $253 billion of revenue. Its price-to-earnings ratio now sits near 34.

Yahoo Finance
Aug 12th, 2026
Nvidia backs open-source AI to boost hardware demand as inference spending hits $23.3B

Nvidia has released its open-source Nemotron 3.5 Lightning model, reinforcing CEO Jensen Huang's public support for open-source artificial intelligence. The move positions the chip maker opposite closed-model proponents like OpenAI and Anthropic in the AI development debate. The strategy serves Nvidia's commercial interests by directing enterprise spending towards hardware rather than expensive proprietary software. "What Nvidia is doing is removing that cost, which means [enterprise clients] have more money for hardware," said Bill Wong, AI research fellow at Info-Tech Research Group. Global AI inference spending is projected to reach $23.3 billion this year, surpassing training expenditure for the first time, according to Gartner. Nvidia holds 90% of the AI training market and has increased its inference market share from 66% to 74% year-on-year.

The Register
Aug 12th, 2026
Nvidia launches router to slash AI costs by 74% using model switching

Nvidia has introduced NeMo Switchyard, a router designed to reduce enterprise AI costs by directing prompts to different models based on cost, latency, or quality requirements. The platform routes requests between expensive proprietary models and cheaper, smaller alternatives, potentially cutting job completion costs by 74% compared to using Claude Opus 4 alone, with approximately six-point accuracy trade-off. Switchyard functions as a proxy between inference servers and models, optimising which model handles each task. Nvidia suggests using smaller, specialised models for simple tasks like generating title cards, whilst reserving larger models for complex work. The company has developed application-specific models, including Nemotron Parse for PDF processing. Similar routing approaches have been adopted by OpenAI and AT&T. The telecommunications giant reportedly saved 80-90% in certain applications by switching to open-weight models, which now power 25% of its AI workloads.

Yahoo Finance
Aug 12th, 2026
Nvidia partners with BlackRock, Goldman Sachs on $500B AI data centre financing deal

Nvidia has partnered with major financial firms including Apollo, BlackRock, Blackstone, and Goldman Sachs to establish a $500 billion facility for AI data centre infrastructure. The initiative will focus on debt financing to provide computing access for Nvidia's largest customers. Nvidia reported quarterly revenue of $81 billion, up 85% year over year, putting its revenue run rate near $330 billion. The company projects $91 billion in revenue for the current quarter. Nvidia maintains financial relationships and ownership stakes in numerous AI companies including Anthropic, OpenAI, Intel, Coreweave, Nokia, Synopsys, and Marvell. The infrastructure partnership is expected to generate additional revenue for Nvidia, supporting its continued top-line growth of over 60% quarterly.