Full-Time

Member of Technical Staff

Inference

Inferact

Inferact

11-50 employees

Open-source LLM inference engine with scalability

Compensation Overview

$200k - $400k/yr

+ Equity

H1B Sponsorship Available

San Francisco, CA, USA

Remote

Remote in the United States possible for exceptional candidates.

Category
Software Engineering (1)
Required Skills
Python
PyTorch
Machine Learning

Get referred to Inferact

See people who can refer or advise you

Requirements
  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
Responsibilities
  • Push the boundaries of what's possible in LLM and diffusion model serving.
  • Optimize how models execute across diverse hardware and architectures.
  • Work at the core of vLLM, optimizing model execution across diverse hardware and architectures.
  • Your work will directly impact how the world runs AI inference.
Desired Qualifications
  • Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
  • Familiarity with RL frameworks and algorithms for LLMs.
  • Experience with multimodal inference (audio,image/video/text).
  • Contributions to open-source ML or system infrastructure projects.
  • Implemented core features in vLLM or other inference engine projects.
  • Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
  • Written widely-shared technical blogs or side projects on vLLM or LLM inference.

Inferact builds AI inference infrastructure by maintaining vLLM, an open-source LLM inference engine, and offering a managed enterprise inference service. vLLM uses PagedAttention to optimize GPU memory, cutting inference costs and latency while preserving model quality. The company supports multiple architectures and hardware, aligns with PyTorch Foundation governance, and pursues open-source collaboration alongside a commercial platform. Its goal is to turn AI inference into a reliable, scalable operating layer of the AI stack, separating model deployment from application development.

Company Size

11-50

Company Stage

Seed

Total Funding

$150M

Headquarters

San Francisco, California

Founded

2025

Get referred to Inferact

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Enterprise demand for cheaper inference directly fits vLLM's core cost-reduction strengths.
  • Broad hardware and model support attracts teams with heterogeneous production stacks.
  • Observability, disaster recovery, and serverless deployment can convert open-source users into paid customers.

What critics are saying

  • NVIDIA and hyperscaler inference products can bundle software, hardware, and support more tightly.
  • A vLLM community fork would weaken Inferact's open-source moat and commercial funnel.
  • Operational failures across many accelerators would damage enterprise trust quickly.

What makes Inferact unique

  • Founded by vLLM maintainers, giving Inferact direct control of the leading open-source inference engine.
  • PagedAttention reduces GPU memory waste, improving throughput, latency, and serving costs.
  • Maintains vendor-neutral open-source governance while commercializing managed inference services.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Growth & Insights

Headcount

6 month growth

12%

1 year growth

12%

2 year growth

12%