Simplify Logo
Designworks Talent

Designworks Talent

AI Training Infrastructure Engineer

Full-Time
No salary listed
Mid
Bellevue, WA, USA
Hybrid

Approximately three days per week in the office. Relocation is encouraged for candidates elsewhere in the U.S.

No H1B Sponsorship

About the job

Requirements
  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.
  • Experience supporting large artificial intelligence models, foundation models, post-training workflows, or similar machine learning systems.
  • Strong understanding of reliability, scalability, and efficiency challenges associated with multi-node graphics processing unit training.
  • Experience integrating training systems with production machine learning pipelines.
  • Strong programming skills and experience working with complex distributed systems.
  • Ability to independently own technically challenging projects in a fast-moving engineering environment.
  • Comfort operating with high ownership and limited process overhead.
Responsibilities
  • Build and scale distributed training infrastructure supporting large artificial intelligence models across large graphics processing unit clusters.
  • Design and improve systems that increase training reliability, efficiency, and resource utilization.
  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.
  • Integrate artificial intelligence models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.
  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.
  • Build tools and automation that improve the developer experience for artificial intelligence researchers and engineers.
  • Establish best practices for training infrastructure, operational processes, and platform reliability.
  • Contribute to the evolution of the artificial intelligence infrastructure platform as an early member of the engineering team.
Desired Qualifications
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.
  • Experience with supervised fine-tuning, reinforcement learning from human feedback, or other post-training workflows.
  • Background operating artificial intelligence training infrastructure at scale within a hyperscaler, artificial intelligence research organization, cloud provider, or graphics processing unit cloud environment.
  • Experience optimizing graphics processing unit utilization, training performance, or distributed system reliability.
  • Familiarity with Kubernetes, containerized artificial intelligence workloads, and large-scale infrastructure platforms.

About the company

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A