Full-Time

MLOps / Cloud Deployment Engineer

Posted on 8/23/2026

Xenon7

Xenon7

No salary listed

Hyderabad, Telangana, India

Hybrid

Hybrid work arrangement in Hyderabad.

Category
DevOps & Infrastructure (2)
,
Required Skills
LLM
Bash
Kubernetes
MLOps
Microsoft Azure
Python
Software Testing
MLflow
Docker
RAG
Role-based Access Control
LangGraph
Terraform
Observability
LangChain
DevOps
Databricks
Snowflake
Requirements
  • At least 5 years of cloud, DevOps, or MLOps engineering experience on AWS, Azure, or Google Cloud Platform.
  • Production deployment experience for machine learning or generative AI systems, including continuous integration and continuous delivery, containerization with Docker and Kubernetes, and infrastructure as code with Terraform.
  • Experience with MLOps tooling such as MLflow, SageMaker Pipelines, or Azure Machine Learning Pipelines, or equivalent tools.
  • Operational experience with large language model or generative AI systems, including observability, cost monitoring, latency optimization, and prompt and model versioning.
  • Hands-on experience with at least one cloud-native AI platform: AWS Bedrock, SageMaker, Azure AI Foundry, Azure OpenAI, or Vertex AI.
  • Strong Python, Bash, and infrastructure scripting skills.
  • Experience with security and governance in regulated environments, including role-based access control, secrets management, auditing, and compliance.
Responsibilities
  • Own continuous integration and continuous delivery pipelines for machine learning models, retrieval-augmented generation applications, and agentic AI systems from experimentation through production.
  • Deploy and operate AI workloads on cloud-native machine learning and artificial intelligence platforms.
  • Build and maintain observability, tracing, and monitoring for large language model and agentic systems, including latency, cost, hallucination rates, tool-call success, and drift detection.
  • Implement model governance and guardrails, including approval gates, kill switches, escalation paths, and audit trails.
  • Manage infrastructure as code using Terraform, Bicep, or equivalent tools for reproducible AI and machine learning environments.
  • Design cost and performance optimization strategies, including token usage tracking, caching, model routing, autoscaling, and warehouse or cluster right-sizing.
  • Own the security posture, including role-based access control, secret management with Key Vault or Secrets Manager, prompt-injection risk mitigation, and auditability for regulated pharmaceutical environments.
  • Partner with data engineers, AI engineers, and Finance stakeholders to move systems from prototype to reliable production.
  • Implement production evaluation frameworks for AI systems, including regression testing, adversarial testing, accuracy tracking, and hallucination monitoring.
Desired Qualifications
  • Experience in the pharmaceutical, life sciences, or regulated financial services domain.
  • Experience operating agentic AI systems in production, including multi-agent orchestration, tool calling, and human-in-the-loop workflows.
  • Operational experience with LangChain, LangGraph, CrewAI, AutoGen, or Semantic Kernel.
  • Experience with Kubernetes-native machine learning platforms such as Kubeflow or Ray.
  • Operational experience with Snowflake or Databricks, including compute governance and cost management.
  • AWS or Azure Machine Learning Engineer certification, Kubernetes CKA or CKAD certification, or Terraform Associate certification.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A