Full-Time

Member of Technical Staff

AMD GPU Performance Engineering

Inferact

Inferact

11-50 employees

Open-source LLM inference engine with scalability

Compensation Overview

$200k - $400k/yr

+ Equity

H1B Sponsorship Available

San Francisco, CA, USA

Hybrid

The role is based in San Francisco; remote work within the US may be considered for exceptional candidates.

Bachelor's

Category
Software Engineering (1)
Required Skills
PyTorch

Get referred to Inferact

See people who can refer or advise you

Requirements
  • A bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or a similar field.
  • Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.
  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
  • Experience optimizing machine learning kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
  • Strong performance profiling and benchmarking skills, including the use of measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.
Responsibilities
  • Build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling.
  • Improve performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations.
  • Make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Desired Qualifications
  • Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other large language model inference systems.
  • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
  • Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel domain-specific languages and backend libraries.
  • Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.
  • Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source machine learning infrastructure.
  • Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
  • Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.

Inferact builds AI inference infrastructure by maintaining vLLM, an open-source LLM inference engine, and offering a managed enterprise inference service. vLLM uses PagedAttention to optimize GPU memory, cutting inference costs and latency while preserving model quality. The company supports multiple architectures and hardware, aligns with PyTorch Foundation governance, and pursues open-source collaboration alongside a commercial platform. Its goal is to turn AI inference into a reliable, scalable operating layer of the AI stack, separating model deployment from application development.

Company Size

11-50

Company Stage

Seed

Total Funding

$150M

Headquarters

San Francisco, California

Founded

2025

Get referred to Inferact

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Inferact raised $150 million seed at $800 million in January 2026.
  • PyTorch Foundation said vLLM reached stable bi-weekly releases and a Q3 2026 roadmap.
  • Inferact's website says it exists to supercharge vLLM adoption, signaling clear demand.

What critics are saying

  • NVIDIA TensorRT-LLM and llm-d compress pricing before Inferact monetizes enterprise inference.
  • vLLM remains foundation-hosted, limiting Inferact control over roadmap, branding, and licensing.
  • If contributors fork vLLM around Inferact features, the company loses its existential moat.

What makes Inferact unique

  • Inferact owns the vLLM creator team behind the dominant open-source inference engine.
  • vLLM sits in PyTorch Foundation governance, giving Inferact neutral credibility and community reach.
  • PagedAttention and broad hardware support make vLLM the default production serving layer.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Growth & Insights and Company News

Headcount

6 month growth

12%

1 year growth

12%

2 year growth

12%
TechCrunch
Jan 22nd, 2026
Inference startup Inferact lands $150M to commercialize vLLM | TechCrunch

The seed round values the newly formed startup at $800 million.