Full-Time

Machine Learning Infrastructure Tech Lead

Updated on 8/1/2026

Reducto

Reducto

51-200 employees

Ingests data for LLMs and RAG

Compensation Overview

$200k - $300k/yr

San Francisco, CA, USA

In Person

Category
AI & Machine Learning (1)
Software Engineering (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Python
Distributed Systems
CUDA
PyTorch
Machine Learning
Observability

Get referred to Reducto

See people who can refer or advise you

Requirements
  • At least 5 years of experience building production infrastructure, including significant machine learning systems experience.
  • Experience leading complex technical projects from an ambiguous problem through production deployment.
  • Strong Python and systems-engineering skills.
  • Understanding of the performance characteristics of modern GPU training or inference workloads.
  • Comfort with Kubernetes and distributed training or serving frameworks.
  • Ability to reason across low-level model performance and higher-level platform architecture.
  • Ability to maintain high standards for quality, precision, and operational reliability.
  • Ability to operate in a fast-changing, high-growth environment.
  • Ability to take full ownership from strategy through execution.
Responsibilities
  • Own the technical direction and roadmap for the machine learning infrastructure.
  • Build and maintain the training and inference stack, balancing fast experimentation with high-performance production serving.
  • Optimize model serving at every layer, including kernels, runtimes, batching, scheduling, and distributed inference.
  • Design systems for reliable multi-node, multi-GPU training and inference.
  • Improve GPU utilization, latency, throughput, reliability, observability, and cost efficiency.
  • Develop benchmarks that identify bottlenecks and guide infrastructure investments.
  • Evaluate state-of-the-art advances in training and inference and apply relevant advances.
  • Build tooling and abstractions that help machine learning engineers move quickly from experiments to production.
  • Partner with machine learning and Platform engineers on architecture, capacity planning, and technical prioritization.
  • Raise the engineering bar through design reviews, mentorship, and hands-on technical leadership.
Desired Qualifications
  • Experience optimizing or implementing CUDA, Triton, or custom model-serving kernels.
  • Meaningful contributions to frameworks such as vLLM, SGLang, PyTorch, TensorRT-LLM, Ray, or related open-source systems.
  • Experience operating distributed inference or training across hundreds or thousands of GPUs.
  • Experience building observability, scheduling, or capacity-management systems for GPU workloads.
  • Experience at an early-stage or high-growth startup.
  • Connecting technical excellence to measurable business impact.

Reducto.ai helps large organizations handle big volumes of data by ingesting, parsing, and chunking documents so that information is easy to retrieve with large language models. Its system breaks down complex documents into meaningful chunks and extracts structured data, making it easier to feed relevant content into retrieval-augmented generation workflows that work with any vector database. The product works by processing pages, applying layout-based chunking and data extraction, and delivering organized content ready for LLM queries, with options for automatic feature parsing as an add-on. The company differentiates itself by offering enterprise-grade, scalable data processing with dedicated compute resources and tiered subscription plans based on page volumes, plus value-added features for large workloads. The goal is to help businesses improve RAG performance and decision-making by turning vast document collections into searchable, usable data.

Company Size

51-200

Company Stage

Series B

Total Funding

$108M

Headquarters

San Francisco, California

Founded

2023

Get referred to Reducto

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Usage surged more than 6x after the Series A.
  • A flexible startup pricing tier broadens adoption beyond enterprises.
  • Opennote acquisition expands into document agents and workflow integration.

What critics are saying

  • Competitors bundle governance, UI, and self-hosting into broader platforms.
  • Foundation-model vendors can embed parsing and structured extraction directly.
  • Customer concentration exposes revenue and credibility to single-account losses.

What makes Reducto unique

  • Reducto combines OCR and vision-language models for complex document understanding.
  • It targets production-grade ingestion, extraction, editing, and workflow orchestration.
  • Customers include Harvey, Rogo, Scale AI, and Fortune 10 enterprises.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Unlimited Paid Time Off

Wellness Program

Parental Leave

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

1%

2 year growth

0%
FinSMEs
Oct 14th, 2025
Reducto Raises $75M in Series B Funding

Reducto raises $75M in Series B funding. Reducto, a San Francisco, CA-based AI document intelligence platform, raised $75M in Series B funding round. The round, which brought Reducto's total funding to date to $108M. was led by Andreessen Horowitz, with participation from existing investors Benchmark, First Round Capital, BoxGroup, and YCombinator. The company intends to use the funds to accelerate development across model research and product capabilities, and scale adoption across both enterprise and the next generation of AI teams. Led by Adit Abraham, co-founder and CEO, and Raunak Chowdhuri, co-founder and CTO, Reducto is a solution for turning complex documents into AI-ready inputs. Since its founding two years ago, the company has advanced a new standard for document understanding by combining traditional optical character recognition (OCR) with modern Vision-Language Models (VLMs), enabling systems to read documents as a human would. Customers range from AI-native startups, including Harvey, Rogo, and Scale AI, to global financial institutions and Fortune 10 enterprises. These companies use it to handle their most complex and mission-critical document workflow, such as, converting pdfs with redlines to text in legal workflows, extracting complex charts for financial due diligence, or high-stakes figure extraction for healthcare decisions.

The Information
Oct 14th, 2025
Reducto AI Secures New Funding Round

Reducto, a startup integrating OCR with advanced AI to interpret documents, has secured new funding. This investment comes just six months after a previous round, highlighting the company's rapid growth and innovation in document data translation.

Just AI News
Oct 14th, 2025
Reducto Secures $75M Led by Andreessen Horowitz

Reducto secures $75M led by Andreessen Horowitz. * Reducto secured $75M Series B led by Andreessen Horowitz, bringing total funding to $108M in under one year. * The AI document intelligence platform processes nearly one billion pages monthly for Harvey, Rogo, Scale AI, and Fortune 10 enterprises. * Andreessen Horowitz led the round with Benchmark, First Round Capital, BoxGroup, and YCombinator participating as existing investors.

Benzinga
Oct 14th, 2025
Reducto AI Secures $75M Series B Funding

Reducto, an AI document intelligence platform, raised a $75 million Series B round led by Andreessen Horowitz, bringing total funding to $108 million. The company, founded two years ago, combines OCR with Vision-Language Models to enhance document understanding. Reducto's platform processes nearly a billion pages monthly for clients like Scale AI and Fortune 10 enterprises. The new funding will accelerate model research and product development, expanding adoption across enterprises and AI teams.

The American Bazaar
Apr 29th, 2025
Document extraction startup Reducto raises $24.5 million in funding

Reducto has also launched two key improvements - a new agentic OCR framework, which automatically reviews Reducto's outputs, catching mistakes and making corrections through a multi-pass VLM framework, similar to having a human in the loop, and smart cost savings for simpler pages.