Full-Time

Member of Technical Staff

Wafer

Wafer

11-50 employees

Fastest open source LLMs for enterprise, at the lowest cost per token.

Compensation Overview

$200k - $300k/yr

+ Equity + $1,000/month housing stipend

H1B Sponsorship Available

San Francisco, CA, USA

In Person

Category
AI & Machine Learning (1)
Software Engineering
Required Skills
LLM
High Performance Computing (HPC)
Machine Learning

Get referred to Wafer

See people who can refer or advise you

Requirements
  • Experience shipping and operating production systems.
  • Ability to use a profiler, read a trace, and independently diagnose problems.
  • Experience owning a customer relationship and resolving issues in a named account.
  • Ability to explain benchmark methodology to machine learning engineers and defend the results.
  • Ability to work on half-defined problems and return with a scoped solution.
Responsibilities
  • Develop and optimize high-performance computing kernels, and work across inference engine internals and serving infrastructure.
  • Develop AI agents for autonomous inference engineering.
  • Design, deploy, and operate heterogeneous clusters across vendors.
  • Own customer accounts end to end, including determining what to prove, building the solution, operating it in production, and maintaining the customer relationship.
  • Conduct technical evaluations using customers' workloads and explain the results to their engineers.
  • Own production for assigned accounts by diagnosing latency changes and rising error rates, fixing or routing issues, and communicating with customers.
  • Inform the product roadmap based on behavior under real production load.
Desired Qualifications
  • Inference experience is helpful.
  • Ability to learn unfamiliar systems and own them quickly.

At Wafer, we are building AI systems that automatically optimize inference workloads across silicon. The goal is fungible token capacity. Any accelerator optimized toward serving inference most efficiently. Wafer is well funded and serves trillions of tokens a month for mission critical workloads. We serve the highest performance inference to fast-growing AI startups.

Company Size

11-50

Company Stage

Series A

Total Funding

$44.1M

Headquarters

San Francisco, California

Founded

2025

Get referred to Wafer

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Autonomous AI performance engineers dynamically profile runtime bottlenecks and rewrite underlying serving stacks and GPU kernels.
  • Flat-rate "Wafer Pass" subscription abstracts traditional per-token API billing to provide predictable costs for continuous AI coding agents.
  • Hardware-agnostic optimization engine automatically tunes LLM performance across NVIDIA, AMD, and custom enterprise silicon.

What critics are saying

  • A specialized focus on inference requires enterprises to maintain separate providers for model training and raw GPU rentals.
  • A curated serverless catalog offers fewer niche model options compared to full-scale legacy cloud hyperscalers.

What makes Wafer unique

  • Autonomous kernel optimization delivers up to 3x faster open-source LLM inference without manual low-level CUDA engineering.
  • Flat-rate subscription pricing eliminates volatile per-token costs and usage cap anxiety for high-volume AI agent workloads.
  • Multi-hardware compatibility reduces cloud vendor lock-in and protects developers against GPU supply shortages.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Unlimited Paid Time Off

Parental Leave

Daily lunch and dinner

Commuter Benefits

$1K/month housing stipend if you live within walking distance (0.5 miles) of the office

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%
SuperbCrew
Sep 6th, 2026
Wafer raises $40M in Series A funding round.

Wafer raises $40M in Series A funding round. S SuperbCrew is a trusted resource for discovering innovative companies, emerging startups, and the latest technology trends. San Francisco-based Wafer closed a $40M Series A, co-led by Marathon and Chemistry, after rapidly scaling its AI driven inference optimization platform for open source models to $8M ARR with a team of under 10. Wafer raised a $40 million Series A at a valuation of approximately $200 million (or higher), co-led by Marathon and Chemistry. Participating investors included Wing Venture Capital, AMD Ventures, Outset Capital, and existing backers Fifty Years and Y Combinator. A strong group of angels joined, among them Jeff Dean (CEO of DiscoveryLoop), Guillermo Rauch (CEO of Vercel), Andy Fang (CTO of DoorDash), Kyle Vogt, Akshay Kothari (COO of Notion), Matthew Prince (CEO of Cloudflare), and Scott Stephenson (CEO of Deepgram). What is Wafer.AI? Wafer was founded in 2025 by University of Chicago alumni Emilio Andere (CEO; background in ML security research published at NeurIPS and weather modeling at Argonne National Lab) and Steven Arellano (experience in high performance computing and AI infrastructure at Google/Bard, Two Sigma, and Sei Labs). The company emerged from Y Combinator's Summer 2025 batch. It raised a $4 million seed in April 2026 led by Fifty Years (with Liquid2, YC, and early angels including Jeff Dean and OpenAI co-founder Wojciech Zaremba), followed by this Series A roughly five months later. Total funding stands near $44.5 million. The team remains small, around 8-10 people, based in San Francisco and operating fully in-person. Wafer builds AI agents that act as autonomous performance engineers. These agents continuously profile production workloads, diagnose bottlenecks, test changes across the full stack (model variants, inference engines, custom kernels, quantization, caching, scheduling, and hardware), verify correctness and reliability, and deploy improvements. The core thesis is that inference optimization should be continuous and automated rather than a one-time, manual, expert heavy process performed before deployment. The stated mission centers on maximizing "intelligence per watt." The company pivoted from primarily selling optimization tooling/agents to third parties toward operating its own high performance inference cloud focused on open source models. It offers serverless and dedicated endpoints that deliver the cost advantages of open models with latency competitive with (or better than) much smaller or proprietary alternatives. It supports heterogeneous hardware including NVIDIA, AMD, and others. In approximately three to four months after launching its inference cloud, Wafer reached $8 million in annual recurring revenue (ARR), with reports of 30x revenue growth over three months and service of roughly 2 trillion tokens per month at peak scaling phases with a team of about seven. Demand outstripped GPU capacity, prompting the raise. Named customers and users include Vercel and Inworld AI; testimonials and case studies reference Neon Health (lowest observed latency that holds under load), DigitalOcean (significant speedups on AMD for models such as Kimi), and others. Independent checks, such as Vercel's tests on GLM-5.2, have shown roughly 2x throughput versus other serverless providers. Performance claims include running models such as GLM-5.2 on AMD MI355X at ~80% of the throughput of NVIDIA's B200 at less than half the cost, and delivering industry leading "tokens per second" figures (e.g., thousands of tok/s per node on various open models) while preserving accuracy. The product emphasizes continuous adaptation: when traffic patterns shift, models update, or new hardware appears, the system re-optimizes. The round occurred amid rising inference costs, persistent low average GPU utilization (often cited around 20% in production), growing enterprise interest in open source models for cost and customization reasons, and hardware diversification beyond NVIDIA. Wafer's ability to make alternative accelerators competitive strengthens the case for multi vendor strategies and improves economics for high volume open model serving. Speed (latency and throughput under realistic traffic) has emerged as a stronger driver of adoption than pure cost in many accounts. The company turned down acquisition offers from multiple cloud and inference providers, preferring to remain independent and scale. Proceeds are directed at automating more of the optimization loop, expanding capacity to meet demand, and growing the team. Immediate hiring focuses on technical staff (kernels, engines, serving infrastructure, customer workload ownership; compensation in the $200k-$300k base range plus equity), growth/marketing, chief of staff, and go to market roles, all requiring full time, five days a week presence in San Francisco. The $40M Series A at a ~$200M+ valuation represents a roughly 50x step-up from the seed in five months, reflecting exceptional early revenue velocity, technical differentiation, and investor conviction in continuous, agent driven inference optimization as a foundational layer for the next phase of AI infrastructure. AMD Ventures' participation aligns with demonstrated gains on non NVIDIA hardware. The combination of a tiny team, multi million ARR, hardware agnostic performance leadership on open models, and high profile technical and commercial backers positions Wafer as a standout in the crowded AI infrastructure landscape, with capital now available to convert demand into scaled, continuously improving production systems. Please email SuperbCrew your feedback and news tips at hello(at)superbcrew.com

The SaaS News
Sep 2nd, 2026
Wafer raises $40M Series A.

Wafer raises $40M Series A. Wafer raises $40M in a Series A funding round co-led by Marathon and Chemistry to automate and improve its AI inference optimization infrastructure. Updated September 02, 2026 Wafer, an artificial intelligence company, has announced a $40 million Series A funding round to advance its platform for continuous AI inference optimization. Investors. This round was co-led by Marathon and Chemistry, with additional participation from Wing, AMD Ventures, Outset Capital, Fifty Years, and Y Combinator. A group of angel investors also participated, including Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, and Scott Stephenson. Wafer use of funds. Wafer intends to use the capital to automate more of its inference optimization loop and accelerate the development of its inference infrastructure. About Wafer. Wafer develops AI technology designed to continuously optimize AI inference. By analyzing workload traffic patterns and performance constraints, the company provides automated deployment optimizations across models, engines, kernels, and hardware to improve performance per dollar. Funding details. Company: Wafer Raised: $40M Round: Series A Funding Date: September 1, 2026 Lead Investor: Marathon, Chemistry Additional Investors: Wing, AMD Ventures, Outset Capital, Fifty Years, Y Combinator, Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, Scott Stephenson Software Category: Artificial Intelligence Source: https://www.wafer.ai/blog/series-a Updated September 02, 2026

LinkedIn
Sep 1st, 2026
We’re excited to announce that we’ve raised a $40M Series A! Co-led by Marathon Management Partners and Chemistry, with participation from Wing Venture Capital, AMD Ventures, Outset Capital, Fifty… | Wafer

We’re excited to announce that we’ve raised a $40M Series A! Co-led by Marathon Management Partners and Chemistry, with participation from Wing Venture Capital, AMD Ventures, Outset Capital, Fifty Years, and Y Combinator, and our existing investors doubling down on Wafer. We are also joined by an incredible list of angels, including Jeff Dean (CEO, DiscoveryLoop), Guillermo Rauch (CEO, Vercel), Andy Fang (CTO, DoorDash), Kyle Vogt (CEO, Bot ), Akshay Kothari (COO, Notion), Matthew Prince (CEO, Cloudflare), Scott Stephenson (CEO, Deepgram), and more! Most inference optimization today is manual, service-heavy, and done one-time before deployment. Wafer’s vision is AI that optimizes AI. Wafer learns from your workload’s traffic patterns and performance constraints and searches for the optimal deployment across model, engine, kernels, and hardware. This capital helps us accelerate towards automating more of the inference optimization loop, so every deployment has the leverage of an expert inference-performance team continuously finding ways to improve performance per dollar. Thank you to the customers who trusted us with their workloads, the partners who built alongside us, the investors who believed in our mission, and the wafer team members who have worked tirelessly to turn this vision into reality.

Crypto Briefing
Sep 1st, 2026
Wafer raises $40M at $200M valuation, optimises GPU performance with AI agents

Wafer, an AI infrastructure optimisation startup using autonomous agents to improve GPU performance, has raised $40 million at a valuation exceeding $200 million. The company also rejected acquisition offers. This follows a $4 million seed round closed in April 2026. Wafer addresses low GPU utilisation in production environments, which typically runs at around 20%. Its technology deploys AI agents to profile and tune inference workloads across different hardware and model architectures, automating work that would otherwise require specialised engineering teams. The startup focuses on non-Nvidia chipsets, positioning itself as an optimisation layer for the fragmented AI inference market. Wafer's seed round was led by Fifty Years, with participation from Liquid2, Y Combinator, and angel investors including Google's chief scientist Jeff Dean and OpenAI co-founder Wojciech Zaremba.

Wafer
Sep 1st, 2026
Wafer raises $40M Series A to build AI that optimizes AI.

Wafer raises $40M Series A to build AI that optimizes AI. The round was co-led by Marathon and Chemistry to accelerate Wafer's vision of AI that continuously optimizes AI inference. Wafer is excited to announce that Wafer has raised a $40 million Series A, co-led by Marathon and Chemistry, with participation from Wing, AMD Ventures, and Outset Capital. Existing investors Fifty Years and Y Combinator are also doubling down on Wafer. Wafer is joined by an incredible group of angels, including Jeff Dean, CEO of DiscoveryLoop; Guillermo Rauch, CEO of Vercel; Andy Fang, CTO of DoorDash; Kyle Vogt, CEO of Bot; Akshay Kothari, COO of Notion; Matthew Prince, CEO of Cloudflare; Scott Stephenson, CEO of Deepgram; and more. AI that optimizes AI. Most inference optimization today is manual, service-heavy, and performed once before deployment. Wafer believe it should be continuous. Wafer learns from a workload's traffic patterns and performance constraints, then finds the optimal deployment across the model, engine, kernels, and hardware. Its vision is AI that optimizes AI - giving every deployment the leverage of an expert inference-performance team that continually finds new ways to improve performance per dollar. What comes next. This capital will help Wafer automate more of the inference optimization loop and accelerate its work toward inference infrastructure that keeps getting better. Thank you to the customers who trusted Wafer with their workloads, the partners who built alongside Wafer, the investors who believed in its mission, and every member of the Wafer team who has worked tirelessly to turn this vision into reality.