Summer 2026

Inference Optimization Intern

Performance Modeling

Institute of Foundation Models

Institute of Foundation Models

Researches and develops foundation models

No salary listed

Sunnyvale, CA, USA

In Person

On-site in Sunnyvale, California.

Bachelor's

Category
AI & Machine Learning

Get referred to Institute of Foundation Models

See people who can refer or advise you

Requirements
  • Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Responsibilities
  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low-level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.
Desired Qualifications
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.
  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware-software co-design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.
Institute of Foundation Models

Institute of Foundation Models

View

IFM builds and studies foundation models—large AI models designed to learn broadly from diverse data and be used across many tasks. Its work centers on academic research to create open, fast, and practical models that address real-world societal needs, rather than narrow applications. The models are developed by a global team across Abu Dhabi, Paris, and Silicon Valley, with a strong emphasis on openness and collaboration to advance AI science and accessibility. Unlike typical private labs that lock models behind paywalls, IFM aims to provide publicly accessible, efficient models that can be used by researchers and developers to solve real problems. The overarching goal is to push the science of foundation models forward while ensuring their benefits reach society at large.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

United Arab Emirates

Founded

N/A

Get referred to Institute of Foundation Models

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AWS, Cerebras, Nebius, and Hugging Face distribution lowers adoption friction immediately.
  • K2 Horizon's 512K context and mobile-sized 0.9B models expand enterprise use cases.
  • MBZUAI's May 2025 launch and January 2026 K2 Think V2 show sustained output.

What critics are saying

  • Open sourcing K2 Horizon invites DeepSeek, Tencent, and Alibaba clones within months.
  • Enterprise buyers can switch to cheaper hosted models if IFM lacks sticky software.
  • If MBZUAI reprioritizes funding, IFM's Paris and Silicon Valley footprint becomes existential burn.

What makes Institute of Foundation Models unique

  • September 3, 2026 K2 Horizon shipped six Apache 2.0 models from 0.9B to 375B.
  • IFM releases weights, code, data, checkpoints, and logs, not just open weights.
  • Abu Dhabi, Silicon Valley, and Paris talent give IFM global recruiting reach.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

Paid Vacation

Paid Holidays

Parental Leave

Employee Assistance Program

Life Insurance

Disability Insurance

401(k) Plan

Wellness Program

Flexible Work Hours

Remote Work Options

Hybrid Work Options

Stock Options

Company Equity

Company News

PR Newswire
Sep 3rd, 2026
Institute of Foundation Models launches K2 Horizon: largest fully open-source AI model fleet with weights, code, and training data

The Institute of Foundation Models has launched K2 Horizon, a fleet of six fully open-source AI models ranging from 0.9 billion to 375 billion parameters. The models include weights, code, training data, and methodologies, allowing researchers and developers to inspect and reproduce the work. The 0.9B, 3.7B, and 7B models set new benchmarks at their respective scales. The smallest model can run on watches and glasses, whilst the 3.7B and 7B models work on phones. The dense 32B model and sparse 36B model suit local hosting, and the 375B flagship model targets enterprise deployments. Launched by Mohamed bin Zayed University of Artificial Intelligence in May 2025, IFM now operates across Silicon Valley, Paris, and Abu Dhabi. The models are available through Hugging Face, vLLM, and SGLang, with API access through partners including AWS and Cerebras. All models use the Apache 2.0 licence.