Full-Time

Machine Learning Infrastructure Engineer

Institute of Foundation Models

Institute of Foundation Models

Researches and develops foundation models

Compensation Overview

$150k - $450k/yr

Sunnyvale, CA, USA

In Person

Category
DevOps & Infrastructure (1)
Required Skills
Kubernetes
Python
PyTorch

Get referred to Institute of Foundation Models

See people who can refer or advise you

Requirements
  • 5+ years of experience in ML systems, infra, or distributed training
  • Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
  • Strong software engineering fundamentals (Python, systems design, testing)
  • Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
  • Ability to implement algorithms across GPUs/nodes based on mathematical specs
  • Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
  • Experience with large-scale machine learning workloads (strong ML fundamentals)
Responsibilities
  • Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
  • Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.
  • Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
  • Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
  • Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
Desired Qualifications
  • Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
  • Familiarity with performance profiling, kernel fusion, or memory optimization
  • Open-source contributions or published research (MLSys, ICML, NeurIPS)
  • CUDA or Triton kernel experience
  • Experience with large-scale pre-training
  • Experience building custom training pipelines at scale and modifying them for custom needs
  • Deep familiarity with training infrastructure and performance tuning
Institute of Foundation Models

Institute of Foundation Models

View

IFM builds and studies foundation models—large AI models designed to learn broadly from diverse data and be used across many tasks. Its work centers on academic research to create open, fast, and practical models that address real-world societal needs, rather than narrow applications. The models are developed by a global team across Abu Dhabi, Paris, and Silicon Valley, with a strong emphasis on openness and collaboration to advance AI science and accessibility. Unlike typical private labs that lock models behind paywalls, IFM aims to provide publicly accessible, efficient models that can be used by researchers and developers to solve real problems. The overarching goal is to push the science of foundation models forward while ensuring their benefits reach society at large.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

United Arab Emirates

Founded

N/A

Get referred to Institute of Foundation Models

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AWS, Cerebras, Nebius, and Hugging Face distribution lowers adoption friction immediately.
  • K2 Horizon's 512K context and mobile-sized 0.9B models expand enterprise use cases.
  • MBZUAI's May 2025 launch and January 2026 K2 Think V2 show sustained output.

What critics are saying

  • Open sourcing K2 Horizon invites DeepSeek, Tencent, and Alibaba clones within months.
  • Enterprise buyers can switch to cheaper hosted models if IFM lacks sticky software.
  • If MBZUAI reprioritizes funding, IFM's Paris and Silicon Valley footprint becomes existential burn.

What makes Institute of Foundation Models unique

  • September 3, 2026 K2 Horizon shipped six Apache 2.0 models from 0.9B to 375B.
  • IFM releases weights, code, data, checkpoints, and logs, not just open weights.
  • Abu Dhabi, Silicon Valley, and Paris talent give IFM global recruiting reach.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Paid Vacation

Paid Holidays

Parental Leave

Employee Assistance Program

Life Insurance

Disability Insurance

401(k) Plan

Wellness Program

Flexible Work Hours

Remote Work Options

Hybrid Work Options

Stock Options

Company Equity

Company News

PR Newswire
Sep 3rd, 2026
Institute of Foundation Models launches K2 Horizon: largest fully open-source AI model fleet with weights, code, and training data

The Institute of Foundation Models has launched K2 Horizon, a fleet of six fully open-source AI models ranging from 0.9 billion to 375 billion parameters. The models include weights, code, training data, and methodologies, allowing researchers and developers to inspect and reproduce the work. The 0.9B, 3.7B, and 7B models set new benchmarks at their respective scales. The smallest model can run on watches and glasses, whilst the 3.7B and 7B models work on phones. The dense 32B model and sparse 36B model suit local hosting, and the 375B flagship model targets enterprise deployments. Launched by Mohamed bin Zayed University of Artificial Intelligence in May 2025, IFM now operates across Silicon Valley, Paris, and Abu Dhabi. The models are available through Hugging Face, vLLM, and SGLang, with API access through partners including AWS and Cerebras. All models use the Apache 2.0 licence.