Full-Time

Researcher

Post-Training

MakerMaker

MakerMaker

No salary listed

San Francisco, CA, USA

In Person

PhD

Category
AI & Machine Learning (1)
Required Skills
PyTorch
Machine Learning
Data Engineering
Reinforcement Learning
Requirements
  • A strong track record of post-training research, including supervised fine-tuning, reinforcement learning, and reward modeling, at a frontier-model lab or equivalent is required.
  • At least 5 years of hands-on machine learning research experience is required.
  • Experience with large-scale data curation and preference-data pipelines is required.
  • Experience designing evaluation suites for capabilities that are not easily benchmarked is required.
  • Fluency in PyTorch or an equivalent framework and comfort working at distributed-training scale are required.
  • Strong statistical judgment sufficient to identify flawed comparisons is required.
  • Strong written communication is required.
Responsibilities
  • Lead post-training research involving supervised fine-tuning, reinforcement learning from human and AI feedback, reinforcement learning from verifiable rewards, direct preference optimization and successor methods, reward modeling, and preference-data design.
  • Design and curate post-training data, including sourcing, filtering, and quality assessment.
  • Build and maintain evaluation suites that measure relevant model capabilities and avoid over-optimizing to benchmarks.
  • Run rigorous experiments involving controls, ablations, and statistical significance, and clearly document internal findings.
  • Scale data pipelines and the infrastructure required to scale training.
  • Identify and characterize failure modes such as reward hacking, distribution drift, and evaluation saturation, and design experiments to address them.
  • Stay current on post-training research literature and evaluate useful methods for implementation.
  • Set research direction, run experiments, and ship results into production.
  • Partner with data, infrastructure, and engineering teams to make the post-training pipeline reliable and fast.
Desired Qualifications
  • A PhD in machine learning, statistics, computer science, or an adjacent field is preferred.
  • Published research at NeurIPS, ICML, ICLR, COLM, RLC, or comparable venues is preferred.
  • Experience with reward hacking detection, scaling reward models, or reinforcement learning from human feedback infrastructure is preferred.
  • Experience with synthetic data generation is preferred.
  • A background in reinforcement learning mathematics, including policy gradients, importance sampling, and off-policy methods, is preferred.
  • Open-source contributions to post-training infrastructure are preferred.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A