Full-Time

Performance Engineer

On-Device Inference

Updated on 8/1/2026

Sarvam

Sarvam

51-200 employees

Full-stack generative AI platform for enterprises

No salary listed

Bengaluru, Karnataka, India

In Person

Category
AI & Machine Learning (1)
Required Skills
PyTorch

Get referred to Sarvam

See people who can refer or advise you

Requirements
  • At least 3 years of experience working on machine learning systems.
  • Solid experience with PyTorch and ONNX export, including dynamic shapes, control flow, and custom operations.
  • Production quantization experience on at least one real model.
  • Experience with at least two of ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, and LiteRT.
  • Profiling fluency on at least one platform.
Responsibilities
  • Take Sarvam models from research-handoff state to production-ready artifacts on at least two target chipsets: Intel xPU, ARM xPU, Apple xPU, Nvidia GPUs, or AMD GPUs.
  • Own one to two model-and-chipset pairs end-to-end and partner with the consuming application team during integration.
  • Quantize, validate accuracy, benchmark, and document each model-and-chipset pair.
  • Author the deployment workbook for each owned model-and-chipset pair.
  • Embed part-time with consuming teams during integration and debug performance and accuracy issues alongside them.
  • Maintain and extend the team's benchmark harness.
Desired Qualifications
  • Custom operation authoring in any runtime.

Sarvam AI builds a full-stack Generative AI platform and research-informed models for enterprise use in India. It combines training of custom large language models and bespoke enterprise models with an enterprise-grade platform for authoring, deployment, and distribution of AI applications. The product works by offering end-to-end tooling: researchers develop and fine-tune models, then enterprises author, deploy, and manage these models within a scalable platform that supports multilingual and diverse Indian business needs. Sarvam AI differentiates itself through its focus on the Indian market, tailoring AI solutions to linguistic diversity and enterprise requirements, and by pairing model development with deployment infrastructure to deliver cost-effective, robust performance for customer deployments. Its goal is to accelerate the adoption of Generative AI in India by making development, deployment, and distribution of AI applications more efficient and affordable for enterprises.

Company Size

51-200

Company Stage

Series B

Total Funding

$275M

Headquarters

Bengaluru, India

Founded

2023

Get referred to Sarvam

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • HCLTech's $150 million backing accelerates enterprise distribution and credibility.
  • Maruti Suzuki and ICAI validate multilingual, regulated-industry demand.
  • Government sovereign AI mandates create large, defensible India-focused deployments.

What critics are saying

  • OpenAI, Anthropic, and Gemini can commoditize Sarvam's multilingual advantage.
  • GPU shortages and export controls constrain model training and deployment.
  • Government procurement delays slow revenue while capital needs keep rising.

What makes Sarvam unique

  • Full-stack sovereign AI for India, not just model APIs.
  • Built for multilingual voices, documents, and local enterprise workflows.
  • Founded by Vivek Raghavan and Pratyush Kumar in August 2023.

Help us improve and share your feedback! Did you find this helpful?