Full-Time

Model Implementation Engineer

Updated on 7/22/2026

Sciforium

Sciforium

11-50 employees

Serverless AI inference platform

Compensation Overview

$165k - $220k/yr

+ Equity

San Francisco, CA, USA

In Person

On-site in San Francisco, California.

Category
AI & Machine Learning (1)
Required Skills
Python
Neural Networks
Pytorch

Get referred to Sciforium

Find people who can refer or advise you

Requirements
  • At least 3 years of industry or research experience in model implementation or applied machine learning.
  • Master of Science (or higher) in Computer Science, Machine Learning, Electrical Engineering, Applied Mathematics, or a related field.
  • Strong programming skills in Python and experience working with modern ML frameworks.
  • Hands-on experience with JAX and/or PyTorch (JAX strongly preferred).
  • Proven experience maintaining and developing model libraries or reusable ML components.
  • Solid understanding of deep learning architectures across multiple domains (e.g., NLP, vision, speech, generative models).
  • Experience implementing models from research papers and adapting them for real-world usage.
  • Ability to work across teams and collaborate with systems and performance engineering groups
Responsibilities
  • Maintain and evolve a large-scale library of modern machine learning models, including but not limited to LLMs, ASR, TTS, image and video models, and diffusion-based systems.
  • Implement new model architectures and research ideas, ensuring correctness, scalability, and production readiness.
  • Rapidly integrate newly released open-source models to enable day-0 support across the platform.
  • Collaborate closely with GPU kernel and systems teams to optimize model execution and improve overall performance.
  • Benchmark models rigorously and ensure they meet internal performance, latency, and efficiency standards.
  • Contribute to the canonicalization and standardization of model implementations across the library.
  • Develop and maintain internal tooling, testing frameworks, and documentation to support model reliability and reproducibility.
Desired Qualifications
  • Experience with model performance optimization and profiling.
  • Familiarity with low-level performance considerations when running models on GPUs/TPUs.
  • Experience working with large-scale model training or inference systems.
  • Contributions to open-source model repositories or ML frameworks.
  • Experience with JAX-first workflows and advanced features (e.g., pjit, xmap, or custom transformations).

Sciforium provides an AI infrastructure stack with a serverless inference platform, giving access to a mix of open-source and proprietary models through a single API. It runs on its own custom-optimized AMD hardware, delivering models without shared cloud servers to reduce costs, boost performance, and improve data privacy. The company differentiates itself by owning the full stack—hardware, software, and models—and by pursuing byte-native multimodal foundation models. Its goal is to simplify and speed up production-ready AI deployment while lowering costs and complexity for large-scale models.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2024

Get referred to Sciforium

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The AI infrastructure market will reach $101.17B in 2026, driven by GPU backlogs and liquid cooling adoption.
  • Enterprises are shifting to hybrid architectures for elastic training and on-premises inference, favoring scalable stacks like Sciforium's.
  • Custom silicon and optical computing adoption accelerate for energy-efficient AI data processing, aligning with Sciforium's hardware strategy.

What critics are saying

  • AMD's native Instinct MI300 series competes directly with Sciforium's stack, risking customer preference for AMD's cloud partners within 9–15 months.
  • Byte-native models lack proven efficiency at scale; major labs may publish negative benchmarks by 2027, invalidating Sciforium's pipeline control strategy.
  • Exclusive reliance on AMD GPUs limits NVIDIA Tensor Core compatibility, forcing enterprises requiring it to bypass Sciforium for NVIDIA-optimized clouds within 6–10 months.

What makes Sciforium unique

  • Sciforium rebuilt AI serving infrastructure from the ground up, vertically integrated with custom AMD hardware.
  • It offers byte-native multimodal foundation models, shifting from token-based processing to raw byte-level data handling.
  • The platform provides serverless AI inference via a single API, targeting developers and enterprises with multimodal AI apps.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Meal Benefits

Company Equity