Full-Time

Lead Software Engineer

Model Serving Platform

Sciforium

Sciforium

11-50 employees

Serverless AI inference platform

Compensation Overview

$230k - $300k/yr

+ Equity

San Francisco, CA, USA

In Person

Bachelor's

Category
AI & Machine Learning (1)
Required Skills
Kubernetes
MLOps
Python
High Performance Computing (HPC)
CUDA
Machine Learning
Docker
Observability
C/C++

Get referred to Sciforium

See people who can refer or advise you

Requirements
  • A bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • At least 5 years of experience designing and building scalable, reliable backend systems or distributed infrastructure.
  • Strong understanding of large language model inference mechanics, including prefill versus decode, batching, and key-value cache.
  • Experience with Kubernetes, Ray, and containerization.
  • Strong proficiency in C++ and Python.
  • Strong debugging, profiling, and system-level performance optimization skills.
  • Ability to collaborate closely with machine learning researchers and translate model or runtime requirements into production-grade systems.
  • Ability to lead technical discussions, mentor engineers, and drive engineering quality.
  • Comfort working from the office.
Responsibilities
  • Lead the technical direction of the model serving platform, owning architecture decisions and guiding engineering execution.
  • Build core serving components, including execution runtimes, batching, scheduling, and distributed inference systems.
  • Develop high-performance C++ and CUDA/HIP modules, including custom GPU kernels and memory-optimized runtimes.
  • Collaborate with machine learning researchers to productionize new multimodal models and ensure low-latency, scalable inference.
  • Build Python application programming interfaces and services that expose model capabilities to downstream applications.
  • Mentor and support other engineers through code reviews, design discussions, and hands-on technical guidance.
  • Drive performance profiling, benchmarking, and observability across the inference stack.
  • Ensure high reliability and maintainability through testing, monitoring, and engineering best practices.
  • Troubleshoot and resolve complex issues across GPU, runtime, and service layers.
Desired Qualifications
  • Experience with machine learning systems engineering, distributed GPU scheduling, or open-source inference engines such as vLLM, Sglang, or TRT-LLM.
  • Experience building large-scale machine learning or MLOps infrastructure.
  • Proficiency in CUDA or ROCm and experience with GPU profiling tools.
  • Experience at an artificial intelligence or machine learning startup, research lab, or large technology infrastructure or machine learning team.
  • Familiarity with multimodal model architectures, raw-byte models, or efficient inference techniques.
  • Contributions to open-source machine learning or high-performance computing infrastructure.

Sciforium provides an AI infrastructure stack with a serverless inference platform, giving access to a mix of open-source and proprietary models through a single API. It runs on its own custom-optimized AMD hardware, delivering models without shared cloud servers to reduce costs, boost performance, and improve data privacy. The company differentiates itself by owning the full stack—hardware, software, and models—and by pursuing byte-native multimodal foundation models. Its goal is to simplify and speed up production-ready AI deployment while lowering costs and complexity for large-scale models.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2024

Get referred to Sciforium

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • SignalFire still lists Sciforium current as of June 2026.
  • LinkedIn shows aggressive hiring for GPUs, ML, and growth in August 2026.
  • First tranche of GPU capacity targets October 2026, suggesting near-term productization.

What critics are saying

  • October 2026 capacity depends on converting one industrial facility on schedule.
  • Twelve jobs across GPU, data center, and model-serving roles signal heavy burn.
  • No public customers or revenue announcements expose Sciforium to existential platform credibility risk.

What makes Sciforium unique

  • AMD-backed vertical stack ties hardware, serving, and models into one platform.
  • Byte-native multimodal foundation models and inference platform share the same infrastructure.
  • Custom AMD GPU operations target lower TCO than shared-cloud inference competitors.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Meal Benefits

Company Equity

Growth & Insights

Headcount

6 month growth

-11%

1 year growth

-11%

2 year growth

-11%