Full-Time

AI Systems Engineer

AI Model, Training & Inference

Posted on 9/10/2026

Deadline 9/10/27
AMD

AMD

10,001+ employees

Designs and sells processors, GPUs, accelerators

No salary listed

Markham, ON, Canada

Hybrid

Hybrid role in Markham, Canada.

Bachelor's, Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Kubernetes
PyTorch
Machine Learning
Observability
DevOps

Get referred to AMD

See people who can refer or advise you

Requirements
  • Hands-on experience shipping large language models on real hardware.
  • Experience writing GPU kernels that improve production metrics.
  • Experience building systems infrastructure such as orchestration, storage, and monitoring for large-scale GPU clusters.
  • Ability to enable and optimize large-scale model training on AMD Instinct GPU clusters.
  • Ability to debug distributed-training issues including gradient norm explosions, nondeterministic behavior across GPU generations, and compute-communication overlap.
  • Ability to optimize RCCL collective communication patterns across multi-node topologies.
  • Experience writing and optimizing high-performance GPU kernels in HIP, Triton, and MLIR for AMD Instinct architectures.
  • Ability to optimize inference-serving frameworks for AMD GPUs and meet production throughput and latency targets.
  • Experience building production quantization pipelines using FP8, FP6, FP4, GPTQ, and AWQ.
  • Ability to collaborate with silicon architecture and pre-silicon teams on software-stack integration validation.
  • Experience contributing to the open ROCm ecosystem, including SDKs, continuous-integration dashboards, documentation, and developer-cloud enablement.
Responsibilities
  • Enable and optimize large-scale model training, including large language models, vision-language models, and mixture-of-experts architectures, on AMD Instinct GPU clusters.
  • Build and maintain training infrastructure for job orchestration, distributed checkpointing, data-loading pipelines, and storage optimization on multi-thousand-GPU Kubernetes clusters.
  • Debug and resolve training-specific issues involving gradient norm explosions, nondeterministic behavior across GPU generations, and compute-communication overlap in distributed training.
  • Optimize RCCL collective communication patterns, including all-reduce, all-gather, and reduce-scatter, across multi-node topologies.
  • Develop monitoring, alerting, and compliance infrastructure for training-cluster health, data security, and service-level agreement adherence.
  • Design and build validation and testing infrastructure using proxy workloads, synthetic benchmarks, and configurable workload generators to validate platform readiness across AMD Instinct GPU generations.
  • Write and optimize high-performance GPU kernels for general matrix multiplication, attention, quantized matrix multiplication, GPTQ, and AWQ using HIP, Triton, and MLIR.
  • Drive inference enablement on new AMD GPU silicon by running frontier models on each new Instinct generation and creating reproducible guides and reference implementations.
  • Optimize vLLM, SGLang, and TorchServe for AMD GPUs, including batching strategies, key-value-cache management, speculative decoding, and continuous batching.
  • Develop inference-acceleration approaches using bio-inspired algorithms, small-language-model-assisted batching, and custom scheduling strategies.
  • Build quantization pipelines for FP8, FP6, FP4, GPTQ, and AWQ for production model deployment.
  • Collaborate with AMD silicon architecture and pre-silicon teams to provide software feedback and validate software-stack integration for next-generation Instinct GPU designs.
  • Build observability and automated-analysis tooling, including log-analysis pipelines, anomaly detection, performance baselining, regression detection, and diagnostic workflows for large-scale GPU clusters.
  • Contribute to the open ROCm ecosystem and AMD developer experience through software development kits, continuous-integration dashboards, documentation, and developer-cloud enablement.
Desired Qualifications
  • Industry experience shipping production artificial-intelligence or machine-learning infrastructure spanning training and inference.
  • Direct experience enabling frontier models, including GPT-4-class models, end-to-end on AMD Instinct hardware.
  • Experience building anomaly-detection, log-analysis, or observability systems for large-scale distributed GPU infrastructure.
  • Familiarity with AMD Instinct MI-series architectures, including MI300X, MI350X, and MI355X, and the RCCL communication library.
  • Contributions to open-source AI frameworks including PyTorch, vLLM, SGLang, DeepSpeed, and Megatron-LM.
  • Experience designing validation frameworks, proxy benchmarks, or synthetic workload suites for GPU infrastructure at scale.
  • Experience with pre-silicon software validation or hardware-software co-verification workflows.
  • Publications or patents in high-performance computing, machine-learning systems, or GPU-kernel optimization.

AMD is a semiconductor company that designs and sells processors, graphics cards, and accelerators for a range of customers, including data centers, AI developers, and enterprises. Its products include Ryzen CPUs for desktops, EPYC CPUs for data centers, and Radeon GPUs for gaming and professional visualization. AMD also offers the ROCm software stack and Instinct accelerators to boost AI and machine learning performance. The company earns revenue from hardware sales, technology licensing, and software that optimizes hardware performance. AMD differentiates itself through a broad portfolio that combines high-performance CPUs, GPUs, and AI accelerators, along with software ecosystems that optimize compute workloads. Its goal is to provide scalable, efficient computing power for gaming, data centers, and AI workloads while pursuing responsible corporate practices.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

1969

Get referred to AMD

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AMD's Q2 2026 revenue reached $11.54 billion, up 50% year over year.
  • Data Center sales hit $6.72 billion in Q2 2026, more than doubled.
  • Anthropic's first gigawatt starts in first half 2027, validating MI450 demand.

What critics are saying

  • AMD booked $440 million China export-control charges in Q2 2026.
  • Broadcom's July 2026 custom-chip push threatens AMD's merchant GPU model.
  • If Google, Meta, and OpenAI standardize on custom ASICs, AMD loses AI relevance by 2028.

What makes AMD unique

  • AMD's July 22, 2026 Anthropic deal pairs MI450 GPUs with Helios rack-scale systems.
  • AMD's August 31, 2026 Saudi JV gives sovereign AI customers local infrastructure.
  • Zen 6 Venice on TSMC 2nm makes AMD first HPC server chip there.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Hybrid Work Options

Growth & Insights and Company News

Headcount

6 month growth

1%

1 year growth

3%

2 year growth

-2%
Yahoo Finance
Sep 8th, 2026
AMD surges 6.7% as Amazon signals $60B AI chip deal with Qualcomm

Advanced Micro Devices rallied approximately 6.7% to $509.635 on Tuesday as investors bet on surging demand for AI computing infrastructure. The chip designer's second-quarter revenue soared 50% to $11.54 billion, with data-centre sales more than doubling to $6.7 billion. The data-centre division now generates roughly 58% of AMD's total revenue. AMD's stock price stands 82.57% above its $279.15 GF Value, indicating investors are paying a premium for AI-fueled growth. Separately, Qualcomm said Amazon could purchase as much as $60 billion of AI data-centre products through a long-term collaboration. Qualcomm is targeting $15 billion in annual data-centre revenue by 2029.

Yahoo Finance
Sep 7th, 2026
AMD's data centre revenue doubles to $6.72B, yet shares fall 5% on margin concerns and China export charges

AMD reported strong Q2 results with Data Centre revenue more than doubling year-over-year to $6.72 billion and Q3 guidance of roughly $13 billion. Despite this, shares fell approximately 5% due to concerns over product-mix margin pressure and $440 million in charges from China export controls on Instinct MI308 shipments. Gaming revenue declined 31% year-over-year to $779 million. The company's non-GAAP gross margin expanded to 56%, though Data Centre AI products carry margins slightly below the corporate average. Analyst consensus remains bullish, with 41 buy ratings versus 10 hold ratings and zero sell recommendations. The consensus price target stands at $613.84, implying significant upside from current levels. AMD shares are up 113.42% year to date but have traded sideways recently as investors await proof of volume production ramps.

Yahoo Finance
Sep 7th, 2026
Broadcom's $17B AI chip revenue dwarfs AMD's $7B, threatening GPU dominance

AMD's data centre revenue doubled to $7 billion, but Broadcom's AI semiconductor revenue reached $17 billion, growing 221% year-over-year. The scale gap highlights a structural threat to AMD's merchant GPU strategy. Three of AMD's major customers — OpenAI, Anthropic, and Meta — are simultaneously developing custom accelerators with Broadcom. Broadcom claims its co-developed chips outperform GPUs at half the cost, potentially shrinking AMD's addressable market. Broadcom generated $14 billion in free cash flow last quarter and recently announced its 15th consecutive dividend increase. AMD's operating margin stands at 27% compared to Broadcom's 68%. The competitive dynamic centres on AMD's general-purpose Instinct GPUs versus Broadcom's custom silicon approach. If frontier AI labs shift to proprietary chips, AMD's growth assumptions for its $1.4 trillion 2030 total addressable market may need revision.

Yahoo Finance
Sep 7th, 2026
AMD and Intel surge 4%+ as investors widen AI trade beyond Nvidia

AMD gained 4.7% and Intel rose 4.5% on 4 September, while Nvidia added only 0.8%. The single-session divergence suggests investors may be broadening AI exposure beyond Nvidia, though it cannot confirm lasting rotation. AMD's second-quarter revenue hit a record $11.5 billion, up 50% year-over-year, with Data Centre representing 58% of the company. Its portfolio includes MI450 accelerators and sixth-generation EPYC processors. Insider Monkey's database counted 164 hedge funds holding AMD in Q2 2026, up from 134 in Q1. Intel's second-quarter Data Centre and AI revenue was $6.3 billion, up 59% year-over-year. However, negative $8.4 billion adjusted free cash flow raises capital efficiency concerns. The database counted 138 hedge funds holding Intel in Q2 2026, up from 112 in Q1.

Yahoo Finance
Sep 7th, 2026
AMD data center segment to surpass 70% of revenue in 2027, up from 58% last quarter

AMD's data centre segment is on track to surpass 70% of the company's total revenue by 2027, driven by a substantial growth gap between business units. In Q2 2025, data centre products generated 58% of AMD's record $11.5 billion revenue, up from 42% the previous quarter. The data centre division, which includes EPYC server processors and Instinct GPUs for AI computing, grew 107% year-over-year to $6.7 billion. Meanwhile, AMD's other divisions—client processors, gaming chips, and embedded products—grew just 8% combined to $4.8 billion. Management expects data centre sales to accelerate in the second half of the year. If current growth rates continue, with data centre moderating to 90% growth and other segments maintaining 8% growth, the data centre segment would reach approximately $12.8 billion by Q2 2027, representing 71% of total revenue.