Full-Time

AI Infrastructure & Platform Operations Engineer

Remote in the EU

Posted on 8/23/2026

Mirantis

Mirantis

501-1,000 employees

Kubernetes management platform with services

No salary listed

Berlin, Germany

Remote

Category
DevOps & Infrastructure (1)
Required Skills
Kubernetes
Computer Networking
Linux/Unix

Get referred to Mirantis

See people who can refer or advise you

Requirements
  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience working within structured operational and incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work within a shift-based operational environment.
Responsibilities
  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.
  • Participate in incident response, root cause analysis, and operational improvement activities.
  • Contribute to improvements in monitoring, observability, automation, and operational processes.
  • Maintain operational documentation, runbooks, and knowledge articles.
Desired Qualifications
  • NVIDIA GPU infrastructure and accelerated computing platforms.
  • InfiniBand networking and NVIDIA UFM.
  • Kubernetes platform operations.
  • AI infrastructure or HPC environments.
  • Site Reliability Engineering (SRE) or Platform Engineering.
  • Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Large-scale distributed systems and production platforms.

Mirantis provides a Kubernetes management platform and related cloud services. It helps businesses run and scale containerized applications in the cloud by automating deployment, management, and scaling of Kubernetes clusters. The product is available as a free download, while the company earns revenue from professional services such as consulting, training, and support. This mix of software and services sets it apart from competitors that rely solely on software sales. Mirantis works with many clients across industries, focusing on helping organizations migrate to modern, scalable cloud solutions. The overall goal is to enable customers to operate their software more efficiently at large scale in the cloud.

Company Size

501-1,000

Company Stage

Acquired

Total Funding

$255M

Headquarters

Mountain View, California

Founded

1999

Get referred to Mirantis

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • IREN announced a $625 million Mirantis acquisition on May 5, 2026.
  • Mirantis launched k0rdent AI Model Registry and Inference Mesh on May 14, 2026.
  • Saturn Cloud partnered with Mirantis on May 12, 2026, reaching over 100,000 developers.

What critics are saying

  • IREN’s May 5, 2026 acquisition still faces closing conditions and regulatory approval.
  • Mirantis competes with Red Hat, SUSE, Broadcom, and VMware across Kubernetes and runtime layers.
  • A bad IREN integration can gut Mirantis’ standalone sales motion and trigger 2027 customer churn.

What makes Mirantis unique

  • Lens Agents, launched April 30, 2026, governs AI agents with policy, audit, and budgets.
  • k0rdent AI spans bare metal, VMs, managed Kubernetes, and sovereign clouds for GPUs.
  • MCR added gVisor on August 5, 2026, delivering stronger container isolation without VMs.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Remote Work Options

Flexible Work Hours

Unlimited Paid Time Off

Paid Vacation

Health Insurance

Dental Insurance

Vision Insurance

Wellness Program

Mental Health Support

Conference Attendance Budget

Professional Development Budget

Stock Options

Company Equity

401(k) Retirement Plan

401(k) Company Match

Phone/Internet Stipend

Home Office Stipend

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

-2%
Mirantis
Aug 5th, 2026
gVisor Support comes to MCR.

gVisor Support comes to MCR. Drew Erny - August 05, 2026 Summarize with AI. This new runtime option provides stronger isolation for container workloads, without moving them into virtual machines. Starting with release 29.6.1, Mirantis Container Runtime (MCR) supports gVisor as an OCI runtime. OCI runtimes are the low-level components that create and run containers. MCR uses runc by default and also supports crun. With gVisor, MCR adds a third supported runtime option for teams that need stronger workload isolation without moving those workloads into full virtual machines. That matters because containers are no longer used only as a convenient packaging format for trusted application code. Increasingly, they are part of the infrastructure layer itself. Service providers use Mirantis software to build public cloud and managed cloud services where customer workloads may arrive from many sources and run on shared infrastructure. Enterprises build private clouds that serve many internal business units, application teams, and development groups, each with different security practices, maturity levels, and operational requirements. For communications service providers, containerized network functions and cloud-native network functions may run on shared infrastructure that handles traffic or control-plane data for many customers. gVisor does not secure those streams inside the workload, but it can help reduce the risk that a compromised container escapes to the shared host or affects neighboring workloads. AI cloud providers and enterprise AI platforms must keep tenant workflows, models, prompts, embeddings, and datasets isolated even when GPUs, storage, and cluster infrastructure are shared. In these environments, stronger runtime isolation is not just a defense-in-depth feature. It is part of the platform trust model. Operators need ways to reduce the risk that a compromised, misconfigured, experimental, or simply unknown container workload can expose the host, affect neighboring workloads, or reach sensitive infrastructure services. gVisor gives MCR users another tool for making those risk-based placement decisions: run workloads that need stronger isolation with gVisor, while continuing to use runc or crun where standard container behavior, compatibility, or performance characteristics are the better fit. gVisor is designed to improve container security by placing an additional isolation layer between the application and the host operating system. It acts as an application kernel: a user-space runtime that implements a Linux-like interface for containerized applications while reducing the host-kernel surface exposed to those applications. This gives operators a middle path between standard container isolation and full VM isolation. That makes gVisor especially useful for environments where containers may run less-trusted, externally supplied, or multi-tenant workloads. In many cases, it can be used as a drop-in replacement for the default runtime, allowing existing container workflows to stay largely unchanged while adding an extra security boundary. As with any sandboxing technology, gVisor involves tradeoffs. Because it intercepts and handles system calls differently from a standard Linux container runtime, syscall-heavy or performance-sensitive workloads may see additional overhead. There are also compatibility considerations: gVisor implements a large portion of the Linux system interface, but not every syscall or kernel feature is fully supported. Most common applications and language runtimes work well, but operators should test workloads with specialized kernel, filesystem, networking, device, or performance requirements before adopting gVisor broadly. MCR can use multiple OCI runtimes on the same machine, so you do not need to choose one runtime for every workload. You can run security-sensitive workloads with gVisor and continue using runc or crun where compatibility or performance characteristics make them a better fit. Get started. For installation and configuration instructions, see the MCR documentation for using gVisor as an alternative container runtime. The docs include the current package names, supported operating systems, repository setup requirements, and runtime configuration steps. The leading, fully-supported secure Container Runtime. If your container applications require exceptional compatibility, reliability, and security, learn more about MCR, the leading fully-supported enterprise container runtime or contact Mirantis.

Goldin Digital Publishing Inc.
Jul 23rd, 2026
Azul Announces Monthly Critical Security Patch Updates for Java Across All Supported LTS Versions

Azul announces monthly Critical Security Patch Updates for Java across all supported LTS versions. July 23, 2026 Azul announced that it will deliver monthly Critical Security Patch Updates (CSPUs) for Java Long-Term Support (LTS) versions for both Azul Core and Azul Prime, starting in August 2026. The traditional quarterly update cadence can no longer keep pace, as a serious vulnerability surfacing just after a scheduled update can sit unpatched for weeks before the next fix ships. Azul is moving to a monthly rhythm to close that exposure window, delivered with the production-grade stability enterprises depend on. Azul's CSPUs will be released monthly, on the third Tuesday of each month, when a high-priority fix is warranted, giving organizations a predictable, plannable security cadence rather than waiting for the next quarterly update. Azul will provide CSPUs across all the LTS versions it supports - Java 8, 11, 17, 21 and 25 - as well as the current release (Java 26). Azul will also deliver CSPUs for the Java 6 and 7 versions it supports, extending the same monthly security cadence to organizations still running older Java versions in production. Azul brings a proven model to this faster cadence. For years, it has delivered Java updates in two forms each quarter: Patch Set Updates (PSUs), which carry the full set of quarterly changes (typically measured in the hundreds), and Critical Patch Updates (CPUs), which deliver security fixes only, built on a stabilized, production-proven code base. Azul's CSPUs extend that same security-only, stability-first CPU model to a monthly rhythm - targeted fixes for identified vulnerabilities tracked as Common Vulnerabilities and Exposures (CVEs), without the unrelated changes that raise regression risk. Azul will continue to work within the OpenJDK community and the OpenJDK Vulnerability Group to advance Java security. "For years, the world's most demanding enterprises have trusted Azul to deliver security and stability together, and on time," said Scott Sellers, co-founder and CEO of Azul. "As AI sharply increases the volume of threats enterprises face, enterprises shouldn't have to choose between the two. Monthly security-only updates are the new standard Azul is setting for how enterprises protect their Java estates." Industry news. July 23, 2026 Azul announced that it will deliver monthly Critical Security Patch Updates (CSPUs) for Java Long-Term Support (LTS) versions for both Azul Core and Azul Prime, starting in August 2026. July 23, 2026 Prismatic announced native large data sync, a set of new capabilities that enables teams to move high volumes of customer data through the same integration platform they already use to build, deploy, monitor, and manage customer-facing integrations. July 23, 2026 Harness and Kong announced an expansion of their strategic partnership to address the growing security challenges posed by AI-driven architectures, autonomous agents, and Model Context Protocol (MCP) deployments. July 22, 2026 Sauce Labs introduced AURA, its AI-Unified Release Assurance platform, to close that gap. July 22, 2026 Harness announced the Harness AppSec Alliance, a strategic partner ecosystem that brings together application security, API and Agent security capabilities on a single platform. July 22, 2026 Mirantis announced that both k0s, its lightweight single-binary Kubernetes distribution, and k0rdent, its open-source distributed container management platform, have achieved CNCF Certified Kubernetes AI Conformance at Kubernetes v1.35. July 21, 2026 Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber to help developers and customers build AI agents in production with higher token efficiency, lower latency, and reliable performance. July 21, 2026 Harness is extending its platform to cover the full AI Agent Development Lifecycle (DLC), giving enterprises a single set of pipelines and controls to build, test, deploy, and run agents the same way they already ship everything else. July 21, 2026 BellSoft announced the general availability of a new hardened builder image for Paketo Buildpacks(TM). July 20, 2026 Testlio unveiled the next evolution of its AI-driven testing platform, LeoCore(TM). July 20, 2026 Agent Island released v1.7.1. July 17, 2026 GitLab(link is external) released GitLab 19.2. As AI generates more code, dependencies, and change than developers can keep up with, GitLab 19.2 brings agentic automation to clear that load. July 16, 2026 Progress Software(link is external) announced that its Progress Chef(link is external) platform now delivers enterprise lifecycle management and configuration capabilities for NVIDIA DGX Spark, enabling IT teams to securely provision, monitor and manage the desktop AI supercomputer at scale. July 16, 2026 Atlassian Corporation announced new capabilities in Jira to advance AI-native software development for every engineering organization. July 16, 2026 Intruder announced the launch of AI Pentesting for web applications, providing on-demand penetration testing.

AiThority
May 15th, 2026
Mirantis brings enterprise-grade controls to AI infrastructure.

Mirantis brings enterprise-grade controls to AI infrastructure. k0rdent adds capabilities enabling enterprises and GPU cloud operators to govern, scale, and monetize sovereign AI services. May 15, 2026 Prev Next 1 of 42,993 Mirantis, delivering Kubernetes-native infrastructure for AI, announced additional capabilities for k0rdent AI, further expanding the platform beyond infrastructure management to help enterprises, neoclouds, and GPU cloud operators monetize AI infrastructure investments. The new k0rdent AI Model Registry and k0rdent AI Inference Mesh enable organizations to securely host, govern, route, and meter AI models and inference services across federated computing resources. Together, the two new products help organizations transform raw GPU infrastructure into governed, revenue-generating AI platforms. Mirantis also introduced k0rdent AI Inference Runtime designed to maximize tokens per GPU-second for improved infrastructure efficiency and utilization. "As organizations move AI projects from experimentation into production, infrastructure teams are increasingly confronting operational and governance challenges around model distribution, inference visibility, compliance enforcement, and GPU economics," said Kevin Kamel, vice president of product development at Mirantis. "Enterprises and GPU operators have largely been forced to stitch together fragile workflows and disconnected tools to operationalize AI. Models cannot be treated the same as containers because they have their own governance, sovereignty, compliance, and lifecycle requirements. The capabilities we're providing today are validated and benchmarked for users." k0rdent AI Model Registry k0rdent AI Model Registry is optimized for AI model storage and distribution workflows. It provides a secure, OCI-native registry for managing large language models (LLMs), fine-tuned variants, quantized builds, and related AI artifacts across distributed infrastructure. The registry reduces the operational complexity often associated with secure AI model distribution. k0rdent AI Inference Mesh k0rdent AI Inference Mesh routes, meters, audits, and enforces policy on every inference request across models, regions, clusters, and providers. It provides a full view of where AI requests are going, what they cost, and any compliance gaps. The new products build on Mirantis' k0rdent AI platform, which focuses on Kubernetes-native AI infrastructure spanning bare metal, virtual machines, managed Kubernetes, and sovereign clouds.

Business Wire
May 14th, 2026
Mirantis launches k0rdent AI Model Registry and Inference Mesh for enterprise GPU monetisation

Mirantis has announced new capabilities for its k0rdent AI platform, expanding beyond infrastructure management to help enterprises and GPU cloud operators monetise AI infrastructure investments. The company introduced k0rdent AI Model Registry for secure model storage and distribution, and k0rdent AI Inference Mesh for routing, metering and policy enforcement across inference requests. The platform addresses operational challenges around model distribution, governance and GPU economics as organisations move AI projects into production. Mirantis also launched k0rdent AI Inference Runtime to maximise tokens per GPU-second for improved efficiency. The new products, available in preview, build on Mirantis' Kubernetes-native AI infrastructure platform spanning bare metal, virtual machines and sovereign clouds. The company serves major enterprises including Adobe, Ericsson and PayPal.

PR Newswire
May 12th, 2026
Saturn Cloud and Mirantis partner to deliver full-stack AI platform for GPU providers and enterprises

Saturn Cloud has partnered with Mirantis to deliver a full-stack AI platform combining infrastructure automation with AI development tools. The partnership integrates Mirantis k0rdent AI with Saturn Cloud's development environment, enabling GPU providers and enterprises to transform bare-metal infrastructure into self-service AI platforms. Mirantis k0rdent AI automates infrastructure provisioning, multi-tenancy, GPU scheduling and lifecycle management across NVIDIA architectures. Saturn Cloud adds the development layer, providing NVIDIA-accelerated environments, distributed training and API endpoints without requiring Kubernetes expertise. The solution targets neoclouds and enterprises seeking production-ready AI platforms whilst maintaining on-premises data security. Saturn Cloud is currently available on Mirantis k0rdent AI infrastructure, with the combined platform supporting over 100,000 developers.

INACTIVE