Full-Time

LLM Dataset Engineer

Updated on 8/7/2026

Sciforium

Sciforium

11-50 employees

Serverless AI inference platform

Compensation Overview

$155k - $210k/yr

+ Equity

San Francisco, CA, USA

In Person

Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
Data Science
Apache Spark
Machine Learning
Computer Vision
Data Analysis

Get referred to Sciforium

See people who can refer or advise you

Requirements
  • At least 5 years of industry experience in Data Science or Machine Learning, including building and managing datasets for foundation models.
  • Expert-level Python proficiency for high-performance code, including multiprocessing, multithreading, and efficient memory management.
  • Experience working with petabyte-scale datasets directly used to train production-grade large language models or large vision models.
  • Experience building massive large language model training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data.
  • Hands-on experience building post-training datasets for Reinforcement Learning from Human Feedback, Direct Preference Optimization, and multi-turn instruction following, including managing human-labeling workflows and quality gold sets.
  • Mastery of data-at-scale frameworks such as Spark and Ray, or high-performance data-loading formats such as WebDataset and Parquet.
Responsibilities
  • Own the end-to-end creation of pre-training datasets for large language models, including defining the mix of web data, code, books, and technical papers to optimize downstream model performance.
  • Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
  • Lead the development of post-training datasets, including supervised fine-tuning instructions, multi-turn dialogues, and preference-modeling data using Reinforcement Learning from Human Feedback and Direct Preference Optimization.
  • Drive the acquisition and processing of vision and video data, including multimodal alignment, video compression, and temporal data consistency.
  • Develop high-throughput data-processing scripts using Python, multiprocessing, and multithreading for large-scale ingestion and transformation.
  • Conduct statistical analysis of training corpora to identify biases, knowledge gaps, and quality regressions.
  • Design pipelines to generate high-reasoning synthetic data using existing models for data labeling and refinement.
Desired Qualifications
  • Experience building large-scale image or video datasets from scratch, including LAION-style pipelines.
  • Familiarity with large-scale crawling of multimodal data and video processing, codecs, and compression.
  • Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
  • A Master's or PhD in a quantitative field focused on data-centric artificial intelligence or information retrieval.

Sciforium provides an AI infrastructure stack with a serverless inference platform, giving access to a mix of open-source and proprietary models through a single API. It runs on its own custom-optimized AMD hardware, delivering models without shared cloud servers to reduce costs, boost performance, and improve data privacy. The company differentiates itself by owning the full stack—hardware, software, and models—and by pursuing byte-native multimodal foundation models. Its goal is to simplify and speed up production-ready AI deployment while lowering costs and complexity for large-scale models.

Company Size

11-50

Company Stage

N/A

Total Funding

N/A

Headquarters

San Francisco, California

Founded

2024

Get referred to Sciforium

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • SignalFire still lists Sciforium as current, and AMD promoted it on 2026-07-15.
  • Jobs boards showed twelve openings in 2026, signaling aggressive hiring and execution.
  • Docs show a live console and API, indicating a real product beyond marketing.

What critics are saying

  • Sciforium remains tiny: Built In listed seven employees on 2026-02-18.
  • Heavy dependence on AMD and seed capital creates concentration risk if priorities shift.
  • A crowded inference market from OpenAI, Anthropic, and Databricks compresses pricing fast.

What makes Sciforium unique

  • Sciforium pairs byte-native multimodal models with a vertically integrated serving stack in 2026.
  • AMD collaboration supports custom hardware optimization, unlike generic cloud inference vendors.
  • Its serverless API unifies text, vision, image generation, and speech-to-text workloads.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

Meal Benefits

Company Equity

Growth & Insights

Headcount

6 month growth

0%

1 year growth

0%

2 year growth

0%