Full-Time

Researcher

Efficient Inference

MakerMaker

MakerMaker

No salary listed

San Francisco, CA, USA

In Person

PhD

Category
AI & Machine Learning (1)
Required Skills
PyTorch
Machine Learning
Data Analysis
Requirements
  • A strong track record of machine learning research on efficiency methods such as quantization, speculative decoding, distillation, mixture-of-experts, sparse attention, or adjacent methods is required.
  • At least 5 years of hands-on research experience is required.
  • Deep familiarity with training and inference performance characteristics is required.
  • Fluency in PyTorch, Jax, or an equivalent framework is required, along with comfort working at the kernel and serving-framework level when methods require it.
  • A track record of moving efficiency research from prototype to production is required.
  • Strong statistical expertise is required to identify flawed comparisons.
  • Published research at NeurIPS, ICML, ICLR, MLSys, or comparable venues is required.
Responsibilities
  • Research and develop quantization methods, including post-training quantization, quantization-aware training, mixed-precision regimes, and low-bit-width arithmetic.
  • Design and evaluate speculative decoding approaches, including draft models, tree attention, parallel speculation, and lookahead decoding.
  • Investigate training-time efficiency methods that compose well with inference, including distillation, sparse attention, mixture-of-experts, low-rank adaptation, and pruning.
  • Run controlled experiments at production scale and characterize performance on real workloads.
  • Co-design methods with the inference engineering team and push results into production.
  • Read the efficient machine learning and efficient inference literature and translate useful ideas into the organization's stack.
  • Publish research when warranted and share findings internally.
  • Partner with model and training researchers so efficiency choices align with model architecture and post-training decisions.
Desired Qualifications
  • A PhD in machine learning, systems, or a related field.
  • Open-source contributions to quantization, speculative-decoding, or efficient-inference libraries.
  • Experience with hardware-aware optimization and accelerator-specific tooling.
  • A background in numerical methods, low-precision arithmetic, or approximate computation.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A