Simplify Logo
uRun

uRun

Stateful inference runtime for real-time AI

ML Performance Engineer - ML Performance

Full-Time
$250k - $395k/yr

+ Equity

Mid
Remote in USA+1 more

More locations: San Francisco, CA, USA

Remote

About the job

Requirements
  • Deep, hands-on CUDA expertise, including production experience writing custom kernels rather than only calling cuBLAS.
  • Strong background in model inference and post-training optimization at scale.
  • Fluency in GPU memory hierarchy, warp scheduling, kernel fusion, and hardware-aware algorithm design.
  • Experience profiling and benchmarking complex inference pipelines to identify and eliminate performance bottlenecks.
  • Ability to operate at the frontier with minimal guidance by identifying problems, designing approaches, and shipping fixes.
Responsibilities
  • Write custom CUDA kernels to unlock performance headroom unavailable through off-the-shelf frameworks.
  • Optimize model inference end-to-end, targeting sub-50-millisecond latency across the inference platform.
  • Drive 10x performance improvements across memory bandwidth, kernel fusion, operator scheduling, and related areas.
  • Implement zero-copy distributed memory optimizations across multi-GPU and multi-node environments.
  • Own GPU utilization and memory management, maximizing hardware floating-point operations per second.
  • Profile, benchmark, and instrument the full inference pipeline to identify and systematically eliminate bottlenecks.
  • Set the performance engineering standard for the team by defining performance targets and building tooling to measure them.
Desired Qualifications
  • Public work in GPU optimization or inference efficiency, such as open-source contributions, a published paper, or a side project demonstrating depth in vLLM, Flash-Attention, TensorRT-LLM, PyTorch, or equivalent.
  • Experience with hardware-aware optimization frameworks such as CuTe, Triton, TileLang, or similar.
  • Familiarity with distributed memory and communication primitives such as NCCL, InfiniBand, NVLink, and RoCE.
  • Contributions to or deep familiarity with PyTorch Distributed, Ray core, or similar systems.
  • Experience optimizing video generation or other high-throughput, latency-sensitive generative workloads.
  • Prior work at an inference-focused company or research lab pushing the boundary of GPU hardware performance.

About the company

uRun builds an infrastructure layer for interactive real-time generative AI by providing a stateful inference runtime. It keeps a per-user session and context loaded in GPU memory, so follow-up prompts are processed instantly without reloading the model. This allows developers and B2B customers to create responsive AI agents, copilots, and tools. The goal is to offer an infrastructure-as-a-service platform that enables real-time, persistent AI interactions at scale.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2025

Get referred to uRun

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • April 2026 launch announced real-time video, avatars, and world model workloads.
  • Blog says LingBot-World Fast shipped from fresh weights in eleven days.
  • August 2026 page lists waitlist interest and early design partners using their own models.

What critics are saying

  • uRun still sits on waitlist, limiting proof of repeatable production demand.
  • Competitors Baseten, Together, Modal, and Fireworks already sell inference infrastructure in 2026.
  • If latency gains fail, hyperscalers and open-source stacks eliminate uRun's niche.

What makes uRun unique

  • April 2026 blog: uRun keeps session state alive between turns at GPU speed.
  • August 2026 AI Engineer page: APIs stream output, persist sessions, and steer generation live.
  • uRun targets video, avatars, and world models, not generic text chat.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

FSA/HSA

Paid Vacation

Company Equity

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%