Full-Time

Research Scientist

Lead Allies

Lead Allies

No salary listed

San Francisco, CA, USA

Hybrid

Hybrid work is listed for San Francisco and the San Francisco Bay Area; remote work is also mentioned in the description.

PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Python
TensorFlow
Neural Networks
PyTorch
Machine Learning
Data Analysis
Requirements
  • Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning large language models with methods such as reinforcement learning from human feedback, direct preference optimization, and contrastive learning.
  • A strong foundation in machine learning and statistics, with experience designing training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment.
  • Experience with dataset design, large-batch training, rigorous evaluation, and ablation studies.
  • A PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field.
  • A strong understanding of large language models and modern deep learning architectures, including Transformers, diffusion models, and reinforcement learning from human feedback.
  • Proficiency in Python and machine learning research libraries such as PyTorch, JAX, or TensorFlow.
  • Demonstrated ability to design and analyze experiments with statistical rigor.
  • Experience publishing research or working on open-source projects in machine learning, Natural Language Processing, or AI evaluation.
  • Comfort working with real-world usage data and designing metrics beyond standard benchmarks.
  • Ability to translate research questions into practical systems and collaborate across engineering and product teams.
  • Passion for open science, reproducibility, and community-driven research.
Responsibilities
  • Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions.
  • Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks.
  • Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences.
  • Collaborate with engineers to implement and scale research findings into production systems.
  • Prototype and test research ideas rapidly while balancing rigor with iteration speed.
  • Author internal reports and external publications that contribute to the broader machine learning research community.
  • Partner with model providers to shape evaluation questions and support responsible model testing.
  • Contribute to the scientific integrity and transparency of the Client leaderboard and tools.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A