M

MakerMaker

Inference Engineer

Full-Time
No salary listed
Mid, Senior
San Francisco, CA, USA
In PersonWork is on-site in San Francisco.

About the job

Requirements
  • At least 3 years of experience building production-grade, large-scale serving infrastructure.
  • Strong distributed systems experience and experience being on-call for important production systems.
  • Fluency in performance profiling and optimization, including reading flame graphs and using analytical, measured approaches to optimization.
  • Experience with GPU-accelerated inference at scale, including multi-GPU, multi-node, batched, and streaming workloads; AMD GPU experience is preferred.
  • Fluency in Python and ability to read and write systems-level code in at least one of C++, CUDA, ROCm, or Triton.
  • A track record of shipping production infrastructure, preferably serving millions of requests across diverse workloads.
  • Good written communication and ability to write a runbook that someone else can follow during an overnight incident.
Responsibilities
  • Build, operate, and harden production inference systems serving large models at high throughput.
  • Own end-to-end performance characteristics, including throughput, latency, cost per token, and reliability under load.
  • Profile real workloads to identify bottlenecks and ship fixes that improve the targeted metric.
  • Implement and integrate inference optimizations from the research team, including quantization, custom kernels, scheduling improvements, and memory management, into production.
  • Design observability into the inference layer through metrics, tracing, and alerting that surface regressions before users notice them.
  • Run capacity planning, autoscaling, and load testing for batch, online, mixed, and agentic workload shapes.
  • Diagnose and resolve production incidents and write postmortems that turn bugs into systemic fixes.
Desired Qualifications
  • Open-source contributions to inference or serving frameworks.
  • Experience with mixed cloud and on-premises deployments.
  • Familiarity with hardware-aware optimization, including memory hierarchy, NCCL/RDMA, and NUMA.
  • Background in compilers, runtimes, or accelerator software stacks.

About the company

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A