Full-Time

GPU Cloud Platform Engineer

Yotta Labs

Yotta Labs

1-10 employees

Decentralized OS orchestrating AI workloads

No salary listed

Remote in USA + 1 more

More locations: Remote in Canada

Remote

Bachelor's, Master's, PhD

Category
DevOps & Infrastructure (1)
Required Skills
Graphics Processing Unit (GPU)
gRPC
Kubernetes
Microsoft Azure
Python
Grafana
Distributed Systems
Software Testing
CUDA
Git
Computer Networking
Docker
AWS
Go
Prometheus
Helm
Google Cloud Platform

Get referred to Yotta Labs

See people who can refer or advise you

Requirements
  • A Bachelor's degree or higher in Computer Science, Software Engineering, Electronic Engineering, or a related field is required.
  • At least 3 years of experience in systems engineering or DevOps is required.
  • At least 5 years of experience in cloud-native development or AI engineering, including at least 2 years of hands-on experience managing and orchestrating Kubernetes multi-cluster environments, is required.
  • Familiarity with the Kubernetes ecosystem and hands-on experience with kubectl, Helm, and multi-cluster deployment, upgrade, scaling, and disaster recovery are required.
  • Proficiency in Docker and containerization technologies, including image management and cross-cluster distribution, is required.
  • Experience with Prometheus and Grafana, including practical experience with GPU fault monitoring and alerting, is required.
  • Hands-on experience with AWS, Google Cloud Platform, or Microsoft Azure and an understanding of cloud-native multi-cluster architecture are required.
  • Familiarity with distributed file systems such as NFS, JuiceFS, CephFS, or Lustre and the ability to diagnose and resolve performance bottlenecks are required.
  • Understanding of high-performance communication protocols such as InfiniBand, RoCE, NVLink, and PCIe is required.
  • Strong communication skills, self-motivation, and team collaboration are required.
Responsibilities
  • Build and operate large-scale, high-performance GPU clusters, ensure stable operation of compute, network, and storage systems, and monitor and troubleshoot online issues.
  • Conduct performance testing and evaluation of multi-node GPU clusters using standard benchmarking tools to identify and resolve performance bottlenecks.
  • Deploy and orchestrate large models, including large language models and video generation models, across multi-cluster environments using Kubernetes; implement elastic scaling and cross-cluster load balancing for efficient service response under high concurrency.
  • Participate in the design, development, and iteration of GPU cluster scheduling and optimization systems; define and lead Kubernetes multi-cluster configuration standards; optimize scheduling strategies such as node affinity and taints/tolerations to improve GPU resource utilization.
  • Build a unified multi-cluster management and monitoring system for cross-region resource monitoring, traffic scheduling, and fault failover; collect GPU memory usage, QPS, and response latency metrics in real time and configure alert mechanisms.
  • Coordinate with IDC providers to plan and deploy large-scale GPU clusters, networks, and storage infrastructure supporting internal cloud platforms and external customer needs.
Desired Qualifications
  • Experience developing and operating model-as-a-service platforms or large-scale model inference clusters, with a proven track record leading multi-cluster system development or performance optimization projects.
  • Proficiency in CUDA programming and the NCCL communication library, with an understanding of high-performance GPUs such as H100.
  • Ability to develop standardized inference APIs using RESTful or gRPC interfaces and automation tools using Golang or Python.
  • Hands-on experience with model quantization, static compilation, and multi-GPU parallelism, including profiling multi-cluster inference processes and identifying memory fragmentation and low compute efficiency.
  • Active engagement with open-source communities such as Hugging Face and GitHub, and understanding of inference frameworks such as Triton, vLLM, and SGLang, including secondary development and optimization for production multi-cluster solutions.
  • Experience with cluster management tools such as Ray, Slurm, KubeSphere, Rancher, or Karmada.

Yotta Labs builds a decentralized operating system for AI workloads, called DeOS, designed to run at planet-scale. It coordinates and optimizes LLM training and inference by scheduling AI tasks across geo-distributed GPUs in a global network. The product works by orchestrating resources and communication between nodes so that AI workloads use available hardware efficiently, enabling very large-scale processing (they describe it as Yottascale). The company differentiates itself by focusing on a decentralized, globally distributed architecture for AI orchestration and high-performance inter-node communication, rather than relying on a single data center. Its goal is to unlock the maximum potential of decentralized AI and push AI workloads toward extremely large-scale processing.

Company Size

1-10

Company Stage

Grant

Total Funding

$300K

Headquarters

Seattle, Washington

Founded

2024

Get referred to Yotta Labs

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The NSF awarded Yotta Labs $296,736 in 2025 for DeAI OS research.
  • Yotta's 2025 Walrus partnership lowers storage overhead for datasets, artifacts, and media outputs.
  • Mysten Labs' Seal and Nautilus integrations extend Yotta's roadmap for permissioned AI data workflows.

What critics are saying

  • Yotta Labs listed five jobs in 2025, signaling a tiny commercialization team.
  • NSF funding ended March 31, 2026; Yotta must convert research into revenue fast.
  • AWS, NVIDIA, and centralized clouds can outspend Yotta before its GPU marketplace scales.

What makes Yotta Labs unique

  • Yotta Labs builds DeOS for AI training across geo-distributed, heterogeneous GPUs.
  • Its NSF-funded scheduler targets up to 80% lower compute costs than centralized clouds.
  • Walrus became Yotta's default storage layer for decentralized AI workflows in 2025.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote Work Options

Flexible Work Hours

Company News

EIN News
Sep 26th, 2025
Yotta Labs Awarded Grant from the U.S. National Science Foundation to Advance Decentralized AI

The project will develop its Decentralized AI Computing OS, designed to enable efficient AI workload on geo-distributed, heterogeneous computing resources.

The Raptor Group
Sep 4th, 2025
Yotta Labs taps Walrus as the dedicated data layer for decentralized AI storage and workflow management

Yotta Labs, builders of the Decentralized Operating System (DeOS) for AI, has partnered with Walrus to deliver high-performance, decentralized storage that reduces overhead for AI workloads and provides seamless, cost-effective data management for Yotta users.