Simplify Logo
SF Tensor

SF Tensor

Hardware-agnostic AI and HPC software stack

Member of Technical Staff - GPU Kernels

Full-Time
$285k - $315k/yr

+ Equity

Mid
San Francisco, CA, USA
In Person

Relocation assistance is offered; work is primarily in the San Francisco office.

About the job

Requirements
  • A track record of hand-writing kernels that match or beat vendor libraries.
  • Comfort reading PTX, SASS, GCN/CDNA ISA, or equivalent machine-level assembly.
  • Fluency with low-level profiling tools such as Nsight Compute, Nsight Systems, rocprof, and omniperf, or equivalent tools.
  • Solid systems programming skills in C++ and CUDA or ROCm/HIP, with a working understanding of how high-level machine learning operations map onto hardware, including where framework layers get in the way.
Responsibilities
  • Write and hand-optimize kernels for real workloads to find the performance ceiling before the search looks for it.
  • Profile at the microarchitectural level, including SM and CU utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput, and occupancy tradeoffs.
  • Debug clock behavior, thermal throttling, and driver paths.
  • Convert kernel knowledge into a machine-searchable structure that the compiler can explore independently.
  • Work below PTX at the ISA level, reasoning about SASS and cubins to emit schedules that PTX cannot express.
  • Build performance models, microbenchmarks, and tooling to predict kernel behavior.
  • Work alongside the formal correctness team so aggressive kernels ship with a formal proof attached.
Desired Qualifications
  • Experience with compiler backends, including MLIR, LLVM, code generation, instruction selection, and scheduling.
  • Research experience in superoptimization, program synthesis, formal verification, or search-based compilation.
  • Experience with silicon beyond NVIDIA, including AMD MI-series, TPU, Trainium, or mobile edge GPUs such as Metal, Mali, and Adreno.
  • Experience with distributed artificial intelligence training.
  • Experience with high-speed interconnects such as NVLink, NVSwitch, InfiniBand, and RoCE.
  • A high-performance computing background involving large-scale scientific computing, MPI, or supercomputing.
  • Experience in driver development or a background in electrical engineering, computer architecture, or hardware design.

About the company

SF Tensor builds an AI/HPC software and infrastructure stack to reduce the infrastructure burden for AI teams. It offers Emma Lang, a hardware-agnostic programming language that runs across GPUs and TPUs without rewriting code, and the SF Tensor Stack, including Kernel Optimizer and Elastic Cloud. The Kernel Optimizer turns models into efficient mathematical forms by simulating hardware topology, often outperforming hand-tuned code, while Elastic Cloud finds cost-effective hardware across clouds and coordinates large-scale training. The goal is to remove vendor lock-in, enable cross-cloud, hardware-agnostic compute, and cut compute costs so AI researchers can focus on innovation.

Company Size

1-10

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2025

Get referred to SF Tensor

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Y Combinator listed SF Tensor in 2025, improving hiring and sales credibility.
  • The company says it can cut compute costs up to 80%.
  • AI labs training frontier models face real pain from cloud portability and GPU inefficiency.

What critics are saying

  • CUDA lock-in remains entrenched; NVIDIA, AMD, and hyperscalers all defend their stacks.
  • Tensor Cloud's private beta signals unfinished product-market fit and revenue risk in 2026.
  • If customers stay on RunPod or Lambda Labs, SF Tensor becomes a niche compiler company.

What makes SF Tensor unique

  • SF Tensor bundles a hardware-agnostic language, kernel optimization, and cross-cloud orchestration.
  • The private beta of Tensor Cloud launched in August 2026.
  • Its pitch targets 1-to-10,000 GPU jobs with automatic cheapest-hardware selection.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Relocation Assistance