Full-Time

GPU Performance / Kernel Engineer

Updated on 9/11/2026

Designworks Talent

Designworks Talent

No salary listed

No H1B Sponsorship

Bellevue, WA, USA

Hybrid

Approximately three days per week in the office. Candidates elsewhere in the U.S. may need to relocate.

Category
Software Engineering (1)
Required Skills
High Performance Computing (HPC)
CUDA
Machine Learning
Requirements
  • Strong experience with GPU kernel development and performance optimization using technologies such as CUDA, ROCm, or comparable GPU programming frameworks.
  • Demonstrated experience improving GPU utilization, reducing latency, or increasing throughput for production artificial intelligence workloads.
  • Strong understanding of GPU architecture, memory hierarchy, parallel computing, and the data path from the application layer to hardware execution.
  • Experience profiling and debugging performance issues in complex artificial intelligence or distributed computing environments.
  • Ability to independently own technically complex problems and drive solutions in a fast-moving engineering environment.
  • Strong systems programming and performance engineering mindset.
  • U.S. work authorization is required.
Responsibilities
  • Profile, analyze, and optimize GPU kernels to improve latency, throughput, and overall utilization.
  • Identify and eliminate data-plane bottlenecks impacting GPU performance across large-scale artificial intelligence workloads.
  • Tune performance-critical workloads across training and inference environments.
  • Work closely with artificial intelligence infrastructure, machine learning, and platform engineering teams to understand workload characteristics and optimize system behavior.
  • Develop benchmarking methodologies and performance measurement practices across GPU infrastructure.
  • Evaluate emerging GPU technologies, performance tools, and optimization techniques as hardware platforms evolve.
  • Contribute to engineering practices that improve GPU efficiency, scalability, and reliability across the fleet.
Desired Qualifications
  • Experience optimizing workloads across multiple GPU platforms, including NVIDIA and AMD architectures.
  • Experience with GPU compiler technologies, runtime optimization, or low-level systems performance.
  • Contributions to open-source GPU performance projects, compiler tooling, or artificial intelligence systems optimization.
  • Background working with large-scale artificial intelligence training, inference platforms, high-performance computing environments, or cloud GPU infrastructure.
  • Familiarity with GPU profiling and optimization tools such as Nsight Systems, Nsight Compute, ROCm profiling tools, or similar technologies.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A