Full-Time

ML Research Engineer

Distributed Training

Metamorphic

Metamorphic

1-10 employees

Tech company offering job opportunities

Compensation Overview

$200k - $280k/yr

+ Equity package

H1B Sponsorship Available

Palo Alto, CA, USA

In Person

Bachelor's

Category
AI & Machine Learning (1)
Required Skills
Distributed Systems
High Performance Computing (HPC)
CUDA
PyTorch
Machine Learning

Get referred to Metamorphic

See people who can refer or advise you

Requirements
  • A bachelor's degree or equivalent experience in Computer Science, Machine Learning, or a related field.
  • Significant software engineering experience with a proven track record of building complex systems.
  • Hands-on experience building and debugging distributed training infrastructure using PyTorch FSDP, DeepSpeed ZeRO, Megatron, TorchTitan, or similar technologies, and optimizing advanced parallelism strategies.
  • Strong understanding of GPU architecture and performance, including memory hierarchy, tensor core utilization, and bandwidth versus compute limitations.
  • Strong understanding of the NVIDIA ecosystem, including CUDA, NCCL, NVLink/NVSwitch topologies, mixed-precision training using MXFP8/NVFP4, and profiling tools.
  • Deep familiarity with PyTorch internals, including torch.distributed, autograd, memory management, and torch.compile.
  • Experience with cloud or high-performance computing environments and job orchestration across hundreds of GPUs.
Responsibilities
  • Build and scale distributed systems that enable training foundation models across thousands of GPUs.
  • Design and optimize the distributed training framework.
  • Implement advanced parallelism strategies.
  • Build fault-tolerant infrastructure.
  • Provide researchers with tooling to run large-scale experiments quickly and reproducibly.
  • Balance research goals with practical engineering constraints.
  • Support collaborative research engineering work, including pair programming and tasks outside the usual job scope when needed.
Desired Qualifications
  • Experience building fault-tolerant training pipelines, including checkpointing, automatic recovery, and infrastructure for reproducible experimentation.
  • Experience with the latest in mixture-of-experts architectures, diffusion model training, or multimodal models.
  • Experience with inference serving frameworks such as vLLM and TensorRT-LLM or building custom inference solutions.

Metamorphic: Current data only provides the company name and a careers page link. No information about its products, services, customers, or market position is available. As a result, a concrete summary answering What does the company do? How do its products work? How is it different from competitors? and What is its goal? cannot be accurately written from the provided information. Please supply additional details or a link to product or about pages so a precise summary can be prepared.

Company Size

1-10

Company Stage

N/A

Total Funding

N/A

Headquarters

New York City, New York

Founded

1994

Get referred to Metamorphic

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Updated August 28, 2026 site and August 14 jobs page show active expansion.
  • Scientific advisory board includes Surya Ganguli, Edward Chang, Jeff Magee, and Amir Zamir.
  • The team sits at the intersection of neuroscience and AI, attracting premium research talent.

What critics are saying

  • Simplify shows no funding, customers, or shipped product, signaling pre-revenue fragility.
  • Hiring only, with no disclosed customer wins, leaves runway dependent on unknown financing.
  • Open-ended frontier NeuroAI research faces talent and execution risk before commercialization.

What makes Metamorphic unique

  • METAMORPHIC builds NeuroAI systems from intelligence itself, not output logs.
  • Founders Andreas Tolias and Sophia Sanborn bring Stanford, MICrONS, and Neuropixels credibility.
  • Palo Alto lab hired technical staff in August 2026 while advertising research-engineering roles.

Help us improve and share your feedback! Did you find this helpful?