Simplify Logo
SF Tensor

SF Tensor

Hardware-agnostic AI and HPC software stack

Member of Technical Staff - GPU Compiler

Full-Time
$285k - $315k/yr

+ Equity

Mid
San Francisco, CA, USA
In Person

Relocation assistance is offered; most work happens at the San Francisco office.

About the job

Requirements
  • Deep experience in compiler infrastructure such as LLVM, MLIR, or similar.
  • A strong background in GPU architecture and low-level optimization using CUDA, ROCm, or similar technologies.
  • Hands-on experience with PTX/SASS, GCN/RDNA assembly, or another GPU instruction set architecture.
  • Familiarity with machine-learning compiler stacks such as XLA, TVM, Triton, torch.compiler, or similar technologies.
  • Solid systems programming skills in C++ and/or Rust.
  • A proven track record of building production-grade compiler infrastructure.
Responsibilities
  • Build and extend MLIR dialects and passes to optimize training and inference workloads.
  • Work in the LLVM backend below ptxas on instruction selection, scheduling, register allocation, and direct cubin emission, and establish equivalent depth for AMD, TPU, and Trainium targets.
  • Expand search-based compiler infrastructure, including agent- and reinforcement-learning-driven program search and formal correctness proofs that make aggressive search safe.
  • Implement classic compiler optimizations tuned for large-scale training.
  • Create hybrid code-generation paths for cases where direct MLIR lowering is not practical.
  • Own testing, benchmarking, and performance-regression systems, including bit-identical hardware models.
  • Work closely with the research team and customer workloads to identify important optimization opportunities.
Desired Qualifications
  • Experience with autotuning or search-based optimization.
  • A background in formal verification, proof assistants, or SMT solvers.
  • Experience writing or maintaining an LLVM backend.
  • A background in distributed systems or multi-device compilation.
  • Contributions to open-source compiler projects.
  • Familiarity with large-scale training infrastructure.
  • Experience with StableHLO or HLO.

About the company

SF Tensor builds an AI/HPC software and infrastructure stack to reduce the infrastructure burden for AI teams. It offers Emma Lang, a hardware-agnostic programming language that runs across GPUs and TPUs without rewriting code, and the SF Tensor Stack, including Kernel Optimizer and Elastic Cloud. The Kernel Optimizer turns models into efficient mathematical forms by simulating hardware topology, often outperforming hand-tuned code, while Elastic Cloud finds cost-effective hardware across clouds and coordinates large-scale training. The goal is to remove vendor lock-in, enable cross-cloud, hardware-agnostic compute, and cut compute costs so AI researchers can focus on innovation.

Company Size

1-10

Company Stage

Seed

Total Funding

$130K

Headquarters

San Francisco, California

Founded

2025

Get referred to SF Tensor

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Y Combinator listed SF Tensor in 2025, improving hiring and sales credibility.
  • The company says it can cut compute costs up to 80%.
  • AI labs training frontier models face real pain from cloud portability and GPU inefficiency.

What critics are saying

  • CUDA lock-in remains entrenched; NVIDIA, AMD, and hyperscalers all defend their stacks.
  • Tensor Cloud's private beta signals unfinished product-market fit and revenue risk in 2026.
  • If customers stay on RunPod or Lambda Labs, SF Tensor becomes a niche compiler company.

What makes SF Tensor unique

  • SF Tensor bundles a hardware-agnostic language, kernel optimization, and cross-cloud orchestration.
  • The private beta of Tensor Cloud launched in August 2026.
  • Its pitch targets 1-to-10,000 GPU jobs with automatic cheapest-hardware selection.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Relocation Assistance