Simplify Logo
SF Tensor

SF Tensor

Hardware-agnostic AI and HPC software stack

Member of Technical Staff - Sandbox Infrastructure

Full-Time
$275k - $315k/yr

+ Equity

Mid
San Francisco, CA, USA
In Person

Relocation assistance is offered; most work happens in the San Francisco office.

About the job

Requirements
  • Strong low-level systems engineering background, including Linux kernel internals, containers, namespaces, cgroups, syscall interception, or hypervisors.
  • Experience in GPU systems engineering, including drivers, runtimes, or scheduling on accelerator fleets.
  • Experience with distributed-systems failure modes, including preemption, partial failure, checkpoint/restore, and handling state that cannot be lost.
  • Proficiency in Go, C/C++, or Rust.
  • A strong bias toward building systems when no vendor-supported solution exists.
Responsibilities
  • Extend the sandboxing stack to new vendors and accelerators.
  • Build and maintain GPU virtualization below the runtime, including gVisor work at the driver and ioctl level.
  • Make sandboxes first-class citizens on spot capacity through preemption-aware scheduling, checkpointing, and rescheduling.
  • Support multi-GPU and multi-node sandboxes, including the required NVLink, NVSwitch, and RDMA paths such as InfiniBand and RoCE.
  • Own live migration end to end, including socket-preserving migration.
  • Guarantee measurement and profiling fidelity as well as sandbox isolation.
  • Work directly with the compiler, post-training, and kernel teams to ensure their throughput is not capped by sandboxes.
Desired Qualifications
  • Direct experience with gVisor, Firecracker, Kata, QEMU/KVM, or similar technologies.
  • Direct experience with CRIU, live migration, or connection-preserving failover work.
  • Familiarity with NCCL/RCCL, RDMA, InfiniBand, or vendor interconnects.
  • Experience running large fleets on spot or other preemptible capacity.
  • Familiarity with bare-metal provisioning, hypervisors, or fleet management at scale.
  • A security background in isolation boundaries and untrusted code execution.

About the company

SF Tensor builds an AI/HPC software and infrastructure stack to reduce the infrastructure burden for AI teams. It offers Emma Lang, a hardware-agnostic programming language that runs across GPUs and TPUs without rewriting code, and the SF Tensor Stack, including Kernel Optimizer and Elastic Cloud. The Kernel Optimizer turns models into efficient mathematical forms by simulating hardware topology, often outperforming hand-tuned code, while Elastic Cloud finds cost-effective hardware across clouds and coordinates large-scale training. The goal is to remove vendor lock-in, enable cross-cloud, hardware-agnostic compute, and cut compute costs so AI researchers can focus on innovation.

Company Size

1-10

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2025

Get referred to SF Tensor

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Y Combinator listed SF Tensor in 2025, improving hiring and sales credibility.
  • The company says it can cut compute costs up to 80%.
  • AI labs training frontier models face real pain from cloud portability and GPU inefficiency.

What critics are saying

  • CUDA lock-in remains entrenched; NVIDIA, AMD, and hyperscalers all defend their stacks.
  • Tensor Cloud's private beta signals unfinished product-market fit and revenue risk in 2026.
  • If customers stay on RunPod or Lambda Labs, SF Tensor becomes a niche compiler company.

What makes SF Tensor unique

  • SF Tensor bundles a hardware-agnostic language, kernel optimization, and cross-cloud orchestration.
  • The private beta of Tensor Cloud launched in August 2026.
  • Its pitch targets 1-to-10,000 GPU jobs with automatic cheapest-hardware selection.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Relocation Assistance