Full-Time

Research Engineer

Updated on 9/11/2026

DatologyAI

DatologyAI

11-50 employees

Automated data curation for GenAI training

Compensation Overview

$180k - $300k/yr

Company Does Not Provide H1B Sponsorship

San Mateo, CA, USA

Hybrid

Four days per week in the office are required.

Category
AI & Machine Learning (1)
Required Skills
Python
PyTorch
Apache Spark
Machine Learning
Data Analysis
Snowflake

Get referred to DatologyAI

See people who can refer or advise you

Requirements
  • At least 3 years of experience building machine learning systems, data infrastructure, or large-scale distributed applications.
  • Strong software engineering fundamentals and fluency in Python.
  • Hands-on experience with PyTorch.
  • Experience with distributed data processing and/or distributed training using tools such as Spark, Ray, Dask, or Snowflake.
  • Experience operating large-scale compute, including GPU clusters and cloud infrastructure.
  • Sufficient machine learning depth to collaborate substantively with researchers.
  • A demonstrated track record of shipping systems that others rely on through production infrastructure, open-source tools, or other artifacts.
Responsibilities
  • Build and scale data processing and curation pipelines operating over massive language, vision, and multimodal datasets, making them fast, reliable, and inexpensive to run.
  • Design experimentation infrastructure that enables scientists to iterate quickly at scale and turn ideas into running experiments within hours.
  • Translate research results into production-grade systems that customers depend on.
  • Profile and optimize large-scale training and data workloads, treating performance and cost as research constraints.

DatologyAI offers automated data curation tools to optimize GenAI training by selecting high-quality, relevant data and removing noisy or harmful data. The core tech analyzes datasets and plugs into existing training pipelines, requiring minimal code changes, and scales from small to petabyte-scale data with usage-based pricing. It differentiates itself with end-to-end automated curation at scale and easy integration, supported by recognized research work and contributions to ImageNet, plus a team with CMU PhD expertise and immigrant-founder VC backing. The goal is to help organizations train better AI models more efficiently and cost-effectively by ensuring high-quality data throughout the training lifecycle.

Company Size

11-50

Company Stage

Series A

Total Funding

$57.7M

Headquarters

Redwood City, California

Founded

2023

Get referred to DatologyAI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • DatologyAI raised $46M in May 2024, giving runway for aggressive hiring.
  • 2026 partnerships with Thomson Reuters and Arcee validate enterprise demand immediately.
  • Thomson Reuters launched Thomson-1.0 using DatologyAI-curated data, proving commercial impact.

What critics are saying

  • OpenAI, Anthropic, and Databricks bundle data tooling, crushing standalone pricing.
  • If Thomson Reuters and Arcee internalize curation, DatologyAI loses repeatable revenue by 2027.
  • A weak benchmark year would kill its claim that curated data beats scaling laws.

What makes DatologyAI unique

  • DatologyAI curates training data, not models, for faster mid-training and post-training.
  • Thomson Reuters used DatologyAI to build a 100B-token legal dataset in 2026.
  • Arcee AI credits DatologyAI with curated 17T public tokens for Trinity-Large.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Unlimited Paid Time Off

Annual Wellness Stipend

Annual Learning and Development Stipend

Relocation Assistance

Company News

SiliconANGLE Media
May 9th, 2024
DatologyAI raises $46M to streamline AI model training data diets

DatologyAI raises $46M to streamline AI model training data diets - SiliconANGLE

DatologyAI
Feb 23rd, 2024
Introducing DatologyAI — Making models better through better data, automatically

Models are what they eat. AI models trained on large-scale datasets have demonstrated jaw-dropping abilities and have the power to transform every aspect of our daily lives, from work to play. This massive leap in capabilities has largely been driven by corresponding increases in the amount of data we train models on, shifting from millions of data points several years ago to billions or trillions of data points today. As a result, these models are a reflection of the data on which they’re train

SiliconANGLE Media
Feb 23rd, 2024
DatologyAI raises $11.65M to automate data curation for more efficient AI training

DatologyAI raises $11.65M to automate data curation for more efficient AI training.

TechCrunch
Feb 22nd, 2024
DatologyAI is building tech to automatically curate AI training datasets | TechCrunch

A new startup, DatologyAI, claims to be able to automatically curate the massive data sets on which AI models train.