Simplify Logo
DatologyAI

DatologyAI

Automated data curation for GenAI training

Research Intern

Summer 2027Posted on 9/19/2026
No salary listed
Internship
San Mateo, CA, USA
Hybrid

Four days in the office per week; the internship runs between May and August 2027.

Company Does Not Provide H1B Sponsorship

About the job

Requirements
  • Strong coding skills.
  • Practical experience and/or publications related to data research, including data pruning and curation, curriculum learning, synthetic data generation, dataset distillation, or the effects of training data on model behavior.
  • Practical experience and/or publications related to embedding models, semantic search, or efficient machine learning.
  • Practical experience and/or publications related to training large vision, language, and multimodal models.
Responsibilities
  • Source, vet, implement, and improve promising ideas from research literature and original research.
  • Conduct novel, high-risk, high-reward research on how data is ingested into future machine learning models.
  • Conduct research guided by concrete customer needs and product improvements.
  • Collaborate closely with engineers, talk to customers, and help shape the product vision.
Desired Qualifications
  • A research topic or area of passion that could improve data curation.

About the company

DatologyAI offers automated data curation tools to optimize GenAI training by selecting high-quality, relevant data and removing noisy or harmful data. The core tech analyzes datasets and plugs into existing training pipelines, requiring minimal code changes, and scales from small to petabyte-scale data with usage-based pricing. It differentiates itself with end-to-end automated curation at scale and easy integration, supported by recognized research work and contributions to ImageNet, plus a team with CMU PhD expertise and immigrant-founder VC backing. The goal is to help organizations train better AI models more efficiently and cost-effectively by ensuring high-quality data throughout the training lifecycle.

Company Size

11-50

Company Stage

Series A

Total Funding

$57.7M

Headquarters

Redwood City, California

Founded

2023

Get referred to DatologyAI

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • August 2026 DataSmith automates data research, reducing labor and widening product breadth.
  • Thomson Reuters launched Thomson on August 24, 2026 using DatologyAI-curated legal data.
  • DatologyAI announced the Deepgram partnership on August 20, 2026 and keeps hiring in 2026.

What critics are saying

  • OpenAI, Anthropic, and cloud vendors can bundle data curation into training stacks by 2027.
  • Thomson Reuters proves customers can internalize DatologyAI's value, compressing pricing and renewal leverage.
  • If model training shifts toward smaller curated datasets, DatologyAI's TAM shrinks and fundraising tightens.

What makes DatologyAI unique

  • DatologyAI curates training data in customers' environments, keeping proprietary corpora inside their cloud.
  • It combines quality filtering, decontamination, distribution balancing, and synthetic data at petabyte scale.
  • Thomson Reuters and Deepgram validations show DatologyAI improves domain models, not just benchmarks.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Unlimited Paid Time Off

Annual Wellness Stipend

Annual Learning and Development Stipend

Relocation Assistance

Company News

DatologyAI
Sep 12th, 2026
DatologyAI raises $46M Series A

DatologyAI is excited to announce that we’ve raised a $46M Series A led by Viv Faga and Astasia Myers from Felicis Ventures, with participation from our existing Seed investors at Radical Ventures and Amplify Partners along with our new investors Elad Gil, M12, and the Amazon Alexa Fund.

SiliconANGLE Media
May 9th, 2024
DatologyAI raises $46M to streamline AI model training data diets

DatologyAI raises $46M to streamline AI model training data diets - SiliconANGLE

DatologyAI
Feb 23rd, 2024
Introducing DatologyAI — Making models better through better data, automatically

Models are what they eat. AI models trained on large-scale datasets have demonstrated jaw-dropping abilities and have the power to transform every aspect of our daily lives, from work to play. This massive leap in capabilities has largely been driven by corresponding increases in the amount of data we train models on, shifting from millions of data points several years ago to billions or trillions of data points today. As a result, these models are a reflection of the data on which they’re train

SiliconANGLE Media
Feb 23rd, 2024
DatologyAI raises $11.65M to automate data curation for more efficient AI training

DatologyAI raises $11.65M to automate data curation for more efficient AI training.

TechCrunch
Feb 22nd, 2024
DatologyAI is building tech to automatically curate AI training datasets | TechCrunch

A new startup, DatologyAI, claims to be able to automatically curate the massive data sets on which AI models train.