Full-Time

GPU Kernel Engineer

Sciforium

Sciforium

11-50 employees

Serverless AI inference platform

Compensation Overview

$190k - $250k/yr

+ Equity

San Francisco, CA, USA

In Person

Bachelor's, Master's, PhD

Category
Software Engineering (1)
Required Skills
LLM
Python
Distributed Systems
CUDA
PyTorch
Machine Learning
C/C++

Get referred to Sciforium

See people who can refer or advise you

Requirements
  • At least 5 years of industry or research experience in GPU kernel development or high-performance computing.
  • A Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
  • Strong programming skills in C++ and Python, with familiarity with machine learning frameworks.
  • Deep expertise in CUDA and ROCm, GPU memory models, and performance optimization strategies.
  • Hands-on experience with Triton and/or JAX Pallas for custom kernel development.
  • Strong understanding of PTX, GPU assembly, and low-level GPU execution.
  • Extensive experience writing and optimizing custom GPU kernels in C++ and PTX.
  • Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks.
  • Experience working with large-scale large language model workloads, including training or inference.
Responsibilities
  • Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.
  • Profile and optimize end-to-end performance of machine learning operations, focusing on large-scale large language model training and inference.
  • Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes.
  • Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate artificial intelligence workloads.
  • Collaborate with machine learning researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack.
  • Work closely with NVIDIA and AMD hardware vendors and stay current on GPU architecture capabilities and compiler and toolchain improvements.
  • Contribute to tooling, documentation, benchmarking suites, and testing frameworks to ensure correctness and performance reproducibility.
Desired Qualifications
  • Experience with AMD GPUs and ROCm optimization.
  • Familiarity with JAX FFI and custom machine learning operator development.
  • Experience with efficient model-serving frameworks such as vLLM and TensorRT.
  • Experience with TPUs, XLA, or similar accelerator programming environments.
  • Contributions to open-source machine learning systems, compilers, or GPU kernels.

Sciforium provides an AI infrastructure stack with a serverless inference platform, giving access to a mix of open-source and proprietary models through a single API. It runs on its own custom-optimized AMD hardware, delivering models without shared cloud servers to reduce costs, boost performance, and improve data privacy. The company differentiates itself by owning the full stack—hardware, software, and models—and by pursuing byte-native multimodal foundation models. Its goal is to simplify and speed up production-ready AI deployment while lowering costs and complexity for large-scale models.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2024

Get referred to Sciforium

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • SignalFire still lists Sciforium as current, and AMD promoted it on 2026-07-15.
  • Jobs boards showed twelve openings in 2026, signaling aggressive hiring and execution.
  • Docs show a live console and API, indicating a real product beyond marketing.

What critics are saying

  • Sciforium remains tiny: Built In listed seven employees on 2026-02-18.
  • Heavy dependence on AMD and seed capital creates concentration risk if priorities shift.
  • A crowded inference market from OpenAI, Anthropic, and Databricks compresses pricing fast.

What makes Sciforium unique

  • Sciforium pairs byte-native multimodal models with a vertically integrated serving stack in 2026.
  • AMD collaboration supports custom hardware optimization, unlike generic cloud inference vendors.
  • Its serverless API unifies text, vision, image generation, and speech-to-text workloads.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Meal Benefits

Company Equity

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%