Full-Time

Founding Machine Learning Infrastructure Engineer

Updated on 8/23/2026

Model AI

Model AI

No salary listed

Palo Alto, CA, USA

In Person

On-site in Palo Alto.

Category
AI & Machine Learning (1)
DevOps & Infrastructure (1)
Required Skills
LLM
Distributed Systems
High Performance Computing (HPC)
CUDA
PyTorch
Machine Learning
Computer Networking
Requirements
  • Strong experience in machine learning systems, distributed systems, or high-performance computing.
  • Experience optimizing inference or training workloads for large models.
  • Familiarity with tensor processing units, graphics processing units, or other accelerators.
  • Experience with one or more of CUDA, Triton, NCCL, JAX/XLA, PyTorch internals, vLLM, SGLang, TensorRT-LLM, distributed inference, or distributed training.
  • Strong systems debugging skills.
  • Comfort working across model code, runtime, infrastructure, and product requirements.
  • High ownership and the ability to operate effectively in an early-stage startup environment.
  • Hands-on technical excellence and strong engineering judgment.
  • End-to-end ownership from design to implementation to production outcomes.
  • Ability to do deep, focused work and sustain execution.
  • Clear communication with teammates, customers, and stakeholders.
  • Comfort with ambiguity, rapid change, and wearing multiple hats.
  • Low ego, high integrity, high accountability, and strong collaboration.
  • Continuous learning and the ability to apply sound judgment in evolving technical environments.
Responsibilities
  • Optimize large-scale large language model inference and serving systems.
  • Improve total tokens per second, decode tokens per second, latency, throughput, and cost efficiency.
  • Work on serving infrastructure for open-source models across different types of accelerators.
  • Improve batching, scheduling, key-value cache management, memory usage, and accelerator utilization.
  • Support long-context inference, including workloads targeting up to one million tokens of context.
  • Debug performance bottlenecks across model execution, runtime, networking, and infrastructure.
  • Use frameworks such as JAX/XLA, PyTorch, vLLM, SGLang, and TensorRT-LLM or related systems.
  • Collaborate with the application team to optimize infrastructure for agentic workloads rather than only generic chatbot inference.
  • Turn research prototypes into reliable, high-performance production systems.
Desired Qualifications
  • Direct tensor processing unit experience is a strong plus.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A