Full-Time

Infrastructure Engineer

AION

AION

11-50 employees

Decentralized GPU marketplace for AI

No salary listed

Bengaluru, Karnataka, India

Hybrid

Category
DevOps & Infrastructure (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Rust
Python
Grafana
High Performance Computing (HPC)
Machine Learning
OpenTelemetry
Infrastructure as Code (IaC)
Role-based Access Control
Go
Prometheus
Terraform
Observability
Grafana Loki
DevOps

Get referred to AION

See people who can refer or advise you

Requirements
  • 6-10 years of experience in infrastructure engineering with a strong focus on observability, monitoring systems, and production service development.
  • Production experience monitoring NVIDIA GPUs using DCGM, NVML, and nvidia-smi, including understanding GPU utilization, memory, temperature, power, ECC errors, and failure modes.
  • Experience deploying and operating Loki, Grafana, Tempo, Mimir, and Prometheus in production, with at least one long-term storage solution such as Thanos, VictoriaMetrics, or Mimir.
  • Experience using Go or Rust for production services, including custom Prometheus exporters, monitoring agents, systemd daemons, and instrumented controllers.
  • Experience building custom Kubernetes controllers and operators using controller-runtime or client-go, with instrumentation and observability for custom resources.
  • Experience designing SLO and SLI frameworks, intelligent alerting strategies, health-check systems, and anomaly detection for distributed workloads.
  • Experience deploying and managing SLURM clusters, understanding job schedulers, and monitoring batch workloads.
  • Advanced Kubernetes expertise, including custom resource definitions, admission controllers, scheduling extensions, and cluster-wide monitoring architectures.
  • Experience monitoring machine-learning training jobs, including GPU efficiency, distributed training metrics, and NCCL performance, as well as inference workloads including latency, throughput, batch processing, and cost per inference.
  • Experience designing metrics architecture, including cardinality management, aggregation strategies, recording rules, retention policies, and balancing observability costs with data granularity at scale.
  • Experience managing observable infrastructure declaratively with GitOps and ArgoCD, building platform abstractions, and creating self-service systems with integrated monitoring.
  • Systems expertise writing unit files, managing service dependencies, deploying monitoring agents as systemd services, and debugging service failures on bare-metal hosts.
  • Experience implementing secure multi-tenant metrics and log isolation, RBAC policies for monitoring data, and platform-wide visibility.
  • Networking proficiency with CNI plugins, network observability, and distributed training communication patterns.
  • Experience using Terraform or similar infrastructure-as-code tools to deploy observable infrastructure and integrate monitoring into provisioning workflows.
Responsibilities
  • Build and deploy comprehensive monitoring for GPU infrastructure using DCGM, NVML, and custom exporters, including metrics pipelines for GPU health, utilization, thermal management, and performance across heterogeneous providers.
  • Deploy and manage production-scale Loki, Grafana, Tempo, Mimir, Prometheus, and Thanos or VictoriaMetrics, including retention strategies, aggregation rules, and scalable query patterns.
  • Write Prometheus exporters in Go or Python for GPU metrics, platform services, and infrastructure components, following metric-naming, labeling, and OpenMetrics standards.
  • Build custom Kubernetes controllers and operators for GPU workload management, scheduling, and resource allocation, with metrics and tracing instrumentation.
  • Design observability for AI training workloads and inference services, covering GPU efficiency, distributed-training performance, resource utilization, latency percentiles, throughput, and cost analytics.
  • Deploy and manage SLURM clusters for high-performance computing workloads, build observability for batch jobs, bridge SLURM and Kubernetes, and create unified monitoring across orchestrators.
  • Write and deploy systemd services for bare-metal GPU nodes, including monitoring agents, metric collectors, and platform daemons, with logging and error handling.
  • Manage infrastructure with ArgoCD and GitOps workflows, create observable platform abstractions, and build declarative self-service capabilities with integrated monitoring.
  • Design actionable alerting for GPU failures, thermal throttling, performance degradation, and workload anomalies; define platform SLOs and monitor reliability.
  • Build secure multi-tenant observability isolation so customers can access only their metrics and logs while operations retain platform-wide visibility, using query-time filtering and RBAC.
  • Implement automated collection of GPU and platform telemetry with OpenTelemetry, managing metric cardinality and storage costs through aggregation.
  • Build monitoring pipelines for GPU utilization, idle time, efficiency, and per-tenant cost allocation, including dashboards for platform economics and provider payout calculations.
  • Design workload-management systems that use observability data to detect hardware failures, trigger workload migration, and handle graceful degradation while surviving infrastructure changes.
  • Build dashboards, APIs, and alerting interfaces that enable providers to monitor hardware contributions and customers to track workload performance in real time.
  • Use observability systems to identify root causes during production incidents involving GPU failures, network bottlenecks, and workload problems; build tooling to reduce MTTD and MTTR.
  • Use observability data to identify distributed-training bottlenecks, inference-latency issues, networking inefficiencies, and opportunities to improve GPU utilization.
Desired Qualifications
  • Founder-level ownership and bias for action.
  • Strategic thinking that connects technical decisions to business impact.
  • Communication and mentoring skills.
  • Ability to thrive in ambiguity, fast-paced environments, and early-stage startup culture.

AION operates a decentralized compute marketplace that lets people and businesses contribute idle GPUs to a shared network and earn passive income. Through this platform, users can buy or sell compute power for AI workloads, making high-performance AI infrastructure more affordable and accessible. Unlike traditional, centralized cloud providers, AION monetizes unused computing resources via a marketplace, offering scalable capacity and democratizing access to AI. The company aims to close the AI wealth gap by enabling widespread ownership of AI compute resources, serving tech-savvy individuals and organizations looking to reduce AI development costs. Revenue comes from facilitating transactions between compute buyers and GPU providers.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

New York City, New York

Founded

2024

Get referred to AION

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AION's July 19, 2026 managed-cloud page shows active commercialization beyond whitepaper promises.
  • The platform targets enterprise bare-metal GPU operators, which shortens supply acquisition and deployment.
  • AION's protocol adds monetization at inference, creating recurring demand from AI agents.

What critics are saying

  • AION's June 2026 Compute Units product resembles securities; SEC scrutiny lands quickly.
  • CoreWeave, AWS, and Nvidia ecosystems already dominate pricing, supply, and enterprise trust.
  • If GPU suppliers defect after weak utilization, AION's marketplace collapses into stranded inventory.

What makes AION unique

  • AION turns idle GPUs into a liquid marketplace, not a traditional cloud stack.
  • AION claims atomic payment-and-verification through x402 headers, bundling usage, proof, and settlement.
  • AION sells compute as a tradable asset, with buybacks and secondary-market exits.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Flexible Work Hours

Wellness Program

Company Equity

Paid Vacation

Growth & Insights

Headcount

6 month growth

-6%

1 year growth

-6%

2 year growth

-6%