Full-Time

AI Software Engineer

AI Inference Platform

ElastixAI

ElastixAI

11-50 employees

Software platform optimizing AI inference costs

No salary listed

Seattle, WA, USA

Hybrid

Three days on-site per week required.

Bachelor's, Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
Graphics Processing Unit (GPU)
Kubernetes
Python
Distributed Systems
CUDA
PyTorch
Docker
REST APIs
C/C++

Get referred to ElastixAI

See people who can refer or advise you

Requirements
  • A Bachelor of Science, Master of Science, or Doctor of Philosophy degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.
  • At least 3 years of professional experience in systems programming, machine learning infrastructure, or distributed inference.
  • Proficiency in C++ and Python, with strong debugging and performance analysis skills.
  • Deep familiarity with one or more large language model serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or DeepSpeed-Inference.
  • Understanding of model deployment internals, including token scheduling, key-value caching, batching, and pipelined inference.
  • Comfort working close to the hardware abstraction layer, including CUDA, PCIe, memory management, or runtime scheduling.
  • Ability to collaborate and communicate cross-functionally in a fast-paced startup environment.
Responsibilities
  • Architect, extend, and optimize core components of the AI serving platform for throughput, latency, and scalability.
  • Customize open-source serving frameworks such as vLLM for proprietary model ingestion and accelerator integration.
  • Develop efficient model partitioning, scheduling, and memory management strategies for multi-device inference.
  • Collaborate with machine learning engineers on model export and runtime optimization, including quantization and graph transforms.
  • Work closely with hardware engineers to influence accelerator interface design and performance tuning.
  • Build application programming interfaces and runtime tools enabling flexible PyTorch-native model deployment on the infrastructure.
  • Profile, debug, and optimize across the full stack, from Python orchestration to C++ kernels and PCIe drivers.
Desired Qualifications
  • Experience with hardware-aware machine learning optimization, compiler/runtime integration, or accelerator software development kits.
  • Hands-on experience profiling graphics processing unit or accelerator workloads.
  • Familiarity with containerized deployments using Docker or Kubernetes.
  • Exposure to distributed systems or large-scale inference clusters.
  • Contributions to open-source machine learning or serving frameworks.

ElastixAI builds a software platform to optimize AI inference for large language models. It reduces cost and complexity by providing a hardware-agnostic, scalable infrastructure that works from edge devices to cloud servers. The platform adapts to different hardware and workloads and can be licensed to chipmakers, cloud providers, and device manufacturers. The goal is to lower the total cost of ownership per token and enable broad, practical deployment of scalable AI inference.

Company Size

11-50

Company Stage

Seed

Total Funding

$34M

Headquarters

Seattle, Washington

Founded

2025

Get referred to ElastixAI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Aug. 2026 press coverage still markets 50x lower TCO, keeping demand interest high.
  • Caplight shows $36M total funding and a Feb. 26, 2026 round, supporting runway.
  • Select enterprise partners and data center operators already access the platform, enabling pilot expansion.

What critics are saying

  • Mid-2026 shipment execution is public; any delay undermines credibility and revenue timing.
  • NVIDIA, hyperscalers, and ASIC startups can crush ElastixAI on performance, ecosystem, and pricing.
  • If customers reject FPGA economics, ElastixAI becomes a niche services company, not a platform.

What makes ElastixAI unique

  • Feb. 2026 launch: FPGA-based inference rack targets GPU server replacement.
  • Mohammad Rastegari and Saman Naderiparizi bring Apple, Xnor, and Waymo systems expertise.
  • Hardware-agnostic software spans hyperscalers, enterprises, and edge deployments.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Paid Parental Leave

Gym Membership

Commuter Benefits

Professional Development Budget

Hybrid Work Options

Paid Holidays

Life Insurance

Paid Holidays

Growth & Insights and Company News

Headcount

6 month growth

-11%

1 year growth

-11%

2 year growth

-11%
Business Wire
Feb 26th, 2026
ElastixAI Emerges From Stealth to Redefine Generative AI Economics via FPGA-Based Supercomputers

ElastixAI Inc. today emerged from stealth to tackle the systemic inefficiencies and high costs of generative AI (GenAI) inference. Founded by former Apple an...

VCNewsDaily
Feb 25th, 2026
ElastixAI raises $18M seed to convert FPGA servers into AI supercomputers

Seattle-based ElastixAI has raised $18 million in seed funding to address inefficiencies and high costs in generative AI inference. Founded by former Apple and Meta machine learning researchers, the company has emerged from stealth with a software platform that converts FPGA-based servers into high-efficiency AI supercomputers. The startup aims to tackle systemic challenges in GenAI inference through its novel technology approach. Details about the funding round's investors were not disclosed.

GeekWire
May 14th, 2025
ElastixAI raises $16M for AI tech

ElastixAI, a Seattle startup founded by former Apple engineers, has raised $16M in a funding round led by FUSE. The company, led by CEO Mohammad Rastegari, is developing an AI inference platform to optimize large language model deployment. Co-founders include Saman Naderiparizi and Mahyar Najibi. ElastixAI aims to improve performance and efficiency for various hardware configurations, offering customizable solutions for hyperscalers and enterprises. Other investors include Catapult, Tyche, Liquid 2 Ventures, and DNX Ventures.