Full-Time

Machine Learning Engineer

Data, Foundational Models

Updated on 8/1/2026

Sarvam

Sarvam

51-200 employees

Full-stack generative AI platform for enterprises

No salary listed

Bengaluru, Karnataka, India

In Person

Category
AI & Machine Learning (1)
Required Skills
Python
Apache Spark
Machine Learning
Data Engineering
Data Analysis

Get referred to Sarvam

See people who can refer or advise you

Requirements
  • A Bachelor of Science or Master of Science degree in Computer Science or a closely related technical field, or equivalent demonstrated experience.
  • At least 3 years of experience building large-scale data systems, including petabyte-scale processing, distributed data pipelines, or comparable work.
  • Hands-on experience with data curation and filtering for large language model training, including the ability to explain a pre-training corpus built end to end and defend its design choices.
  • Deep familiarity with distributed data processing frameworks such as Spark, Ray, Beam, or Dask, or equivalent frameworks, and the storage systems that support them.
  • Strong Python skills and comfort with low-level data-path components such as tokenization, sharding, packing, and input/output patterns, including their performance tradeoffs.
  • Meaningful open-source contributions to the data-tooling ecosystem, such as datasets, deduplication libraries, filtering frameworks, or substantive work on widely used open-data releases.
Responsibilities
  • Design and build large-scale data pipelines for pre-training and post-training at petabyte scale, including ingestion, parsing, normalization, filtering, deduplication, tokenization, and packing.
  • Develop and continually improve quality-filtering systems, including model-based quality classifiers and contamination detection.
  • Own data-mixture design, curriculum, and annealing strategies in partnership with the research team, ensuring the data a model sees, its proportions, and its ordering are precisely documented.
  • Build tooling that enables researchers and engineers to analyze, slice, attribute, and debug data.
  • Scale pipelines to handle multilingual corpora, code, mathematics, multi-source web data, and licensed datasets while tracking provenance and licensing end to end.
  • Partner with the training-infrastructure team to ensure data does not bottleneck production training runs.
Desired Qualifications
  • Direct experience building or working with large open pretraining corpora.
  • Experience with multilingual data collection, normalization, quality scoring, and mixing across many languages.
  • Hands-on experience with model-based data quality classifiers, contamination detection, or data attribution research.
  • Familiarity with tokenization research and the practical implications of tokenizer choices on training.
  • First-author papers or technical reports on data curation, quality, or pretraining mixtures.

Sarvam AI builds a full-stack Generative AI platform and research-informed models for enterprise use in India. It combines training of custom large language models and bespoke enterprise models with an enterprise-grade platform for authoring, deployment, and distribution of AI applications. The product works by offering end-to-end tooling: researchers develop and fine-tune models, then enterprises author, deploy, and manage these models within a scalable platform that supports multilingual and diverse Indian business needs. Sarvam AI differentiates itself through its focus on the Indian market, tailoring AI solutions to linguistic diversity and enterprise requirements, and by pairing model development with deployment infrastructure to deliver cost-effective, robust performance for customer deployments. Its goal is to accelerate the adoption of Generative AI in India by making development, deployment, and distribution of AI applications more efficient and affordable for enterprises.

Company Size

51-200

Company Stage

Series B

Total Funding

$275M

Headquarters

Bengaluru, India

Founded

2023

Get referred to Sarvam

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • HCLTech's $150 million backing accelerates enterprise distribution and credibility.
  • Maruti Suzuki and ICAI validate multilingual, regulated-industry demand.
  • Government sovereign AI mandates create large, defensible India-focused deployments.

What critics are saying

  • OpenAI, Anthropic, and Gemini can commoditize Sarvam's multilingual advantage.
  • GPU shortages and export controls constrain model training and deployment.
  • Government procurement delays slow revenue while capital needs keep rising.

What makes Sarvam unique

  • Full-stack sovereign AI for India, not just model APIs.
  • Built for multilingual voices, documents, and local enterprise workflows.
  • Founded by Vivek Raghavan and Pratyush Kumar in August 2023.

Help us improve and share your feedback! Did you find this helpful?