Full-Time

AI Infrastructure Engineer

GPU

Pragmatike

Pragmatike

Global remote technology talent

No salary listed

Dubai - United Arab Emirates + 2 more

More locations: Remote in Spain | Remote in Italy

Remote

Must work within EMEA time zones.

Category
AI & Machine Learning (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
MLOps
Python
Distributed Systems
CUDA
Machine Learning
MLflow
Infrastructure as Code (IaC)
Terraform
Observability
DevOps
Helm
Requirements
  • Four or more years of experience in MLOps, Platform Engineering, Site Reliability Engineering, or similar infrastructure roles focused on ML systems.
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent.
  • Strong background in container orchestration and operating GPU-based workloads in production.
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines.
  • Proficiency in Python and infrastructure-as-code tools such as Terraform, Helm, or similar.
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering.
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows.
  • Ability to operate independently in a remote-first environment.
Responsibilities
  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent.
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models.
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers.
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance.
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health.
  • Manage model registries and continuous integration/continuous delivery pipelines enabling automated and reproducible model deployments.
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities.
  • Define engineering best practices and contribute to platform scalability.
Desired Qualifications
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI.
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems.
  • Experience with cost optimization across different GPU types and inference workloads.
  • Background in early-stage startups or greenfield infrastructure projects.
  • Proven experience building production systems from scratch rather than maintaining legacy platforms.

Pragmatike connects companies with remote technology professionals across international markets. It recruits for software engineering, data, product, design, infrastructure and related digital roles, helping employers identify specialists outside a single local talent pool. The company combines recruiter-led evaluation with a platform-oriented matching process and supports remote hiring across borders. Its identity is technology talent acquisition and workforce access rather than a software development agency delivering client projects itself. Its workforce brings together technical specialists, client delivery teams and business operations.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A