Internship

Research Fellow

Mechanistic Interpretability

Posted on 7/21/2026

Vmax AI Corp

Vmax AI Corp

No salary listed

San Francisco, CA, USA

Hybrid

Hybrid arrangement possible for exceptional candidates; based in the San Francisco office.

Category
AI & Machine Learning (1)
Required Skills
Python
Neural Networks
Pytorch
Reinforcement Learning
Requirements
  • Currently enrolled in a PhD program in machine learning, computer science, artificial intelligence, computational neuroscience, mathematics, or a related technical field. Exceptional candidates with equivalent research experience may also be considered.
  • Track record of research excellence or strong research promise, demonstrated through publications, preprints, open-source work, technical projects, competitions, or publicly available artifacts.
  • Working understanding of reinforcement learning.
  • Familiarity with mechanistic interpretability, representation analysis, or empirical methods for understanding neural networks.
  • Strong programming ability in Python and experience with at least one major ML framework such as PyTorch or JAX.
  • Clear written and verbal communication of technical ideas.
Responsibilities
  • Develop mechanistic interpretability methods for understanding internal representations, features, circuits, and computations in language models and agents.
  • Investigate how model internals can be used to generate intrinsic rewards, auxiliary objectives, diagnostics, or training signals for reinforcement learning.
  • Design and run experiments that test whether interpretability-derived signals improve learning, exploration, generalization, robustness, or sample efficiency.
  • Compare internally derived rewards against baselines such as human-generated verifiers, reward models, task-level outcome rewards, and standard intrinsic motivation methods.
  • Use techniques such as probing, activation analysis, sparse autoencoders, causal interventions, feature attribution, or representation analysis to study model behavior.
  • Analyze failure modes, including reward hacking, spurious features, non-causal correlations, objective misspecification, and overfitting to narrow evaluation distributions.
  • Build research code, evaluation harnesses, and experimental infrastructure that make results reproducible and useful to the broader team.
  • Communicate research progress clearly through written updates, internal presentations, and final project outputs.
Desired Qualifications
  • Experience with LLM post-training methods
  • Familiarity with intrinsic motivation, unsupervised RL, auxiliary objectives, representation learning for RL, or curiosity-driven learning.
  • Experience with scalable ML experimentation, distributed training, experiment tracking, or reproducible research infrastructure.
  • Interest in turning mechanistic understanding into practical training methods, rather than only analyzing models after training.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A