Full-Time

Postdoctoral Researcher

Jain Lab

Arc Institute

Arc Institute

201-500 employees

Non-profit research institute advancing foundational science

Compensation Overview

$80k/yr

Palo Alto, CA, USA

In Person

PhD

Category
Lab & Research (2)
,
Required Skills
Biochemistry

Get referred to Arc Institute

See people who can refer or advise you

Requirements
  • PhD in metabolism, animal physiology, molecular biology, biochemistry, genomics, or related field
  • Excellent written and verbal communication skills.
  • Demonstrated ability to work in a fast-paced environment and be both an independent thinker and a highly collaborative team player.
Responsibilities
  • Find new functions for enzymes or cofactors (vitamins)
  • Contribute to our molecular understanding of how key metabolites are sensed by the body.
  • Develop novel therapeutic strategies for nutrient-based therapies.
  • Collaborate with post-docs and students to understand how enzymes and metabolites interact for key biochemical functions.
  • Publish, present, and represent that lab in journals and conferences.
  • Present at lab meetings, and participate in Arc-wide activities (seminars, symposiums, etc)

Arc Institute is a non-profit research institution in Palo Alto that pursues curiosity-driven basic science and technology development, aiming to accelerate scientific progress and shorten the path from discovery to patient impact. It collaborates with Stanford University, UCSF, and UC Berkeley, and organizes its work around people rather than specific projects, supporting long-term research agendas. Researchers team across disciplines to study root causes of diseases such as cancer, neurodegeneration, and immune dysfunction, with a focus on understanding disease mechanisms and deploying new technologies at scale to enable practical applications.

Company Size

201-500

Company Stage

N/A

Total Funding

N/A

Headquarters

Palo Alto, California

Founded

2021

Get referred to Arc Institute

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • The Audacious Project added major capital for Arc’s Virtual Cell Initiative in February 2026.
  • BioReason-Pro and Evo 2 broaden Arc’s platform from sequence generation to protein reasoning.
  • Open releases and benchmark leadership attract elite scientists who want visible, publishable impact.

What critics are saying

  • Virtual Cell Challenge results showed pure AI underperformed statistical baselines in July 2026.
  • Arc’s open datasets commoditize its edge, letting Nvidia, Meta, and startups reuse the playbook.
  • Arc depends on donor-backed nonprofit funding; priority shifts could freeze labs by 2028.

What makes Arc Institute unique

  • Arc’s July 2026 Science atlas spans 86,689 cells across 16 human tissues.
  • Arc’s June 2026 Proto turns biological design into programmable, multi-objective workflows.
  • Arc’s open Virtual Cell Atlas exceeded 600 million curated cells by August 2026.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Performance Bonus

Growth & Insights and Company News

Headcount

6 month growth

-1%

1 year growth

0%

2 year growth

0%
San Diego Biotechnology Network
Jul 23rd, 2026
Single-Cell atlas simultaneously maps 3D genome architecture and DNA methylation.

Single-Cell atlas simultaneously maps 3D genome architecture and DNA methylation. July 23, 2026 SDBN news, syndication comments off on single-cell atlas simultaneously maps 3D genome architecture and DNA methylation. Scientists at the Salk Institute and the Arc Institute, along with their collaborators, unveiled the first body-wide single-cell atlas of two major epigenetic systems: three-dimensional genome folding and DNA methylation, measured simultaneously in the same cells. The atlas spans 86,689 cells from 16 human tissues, revealing 35 major cell types and 206 subtypes, and is freely available online. The work is part of the National Institutes of Health's 4D Nucleome (NIH 4DN) program, which aims to understand how the genome is organized in space and time to regulate gene expression in health and disease. Because the two epigenetic layers were measured together, the researchers could compare what each layer says about a cell's identity. And while often the two pictures agree, they found that sometimes they do not. Discover more Biological Sciences

AICell Lab
Jul 18th, 2026
Lab newsletter - july 18, 2026: data decides.

Lab newsletter - july 18, 2026: data decides. AI for life science - daily digest Two big "virtual organ" efforts reported results this stretch, and both point at the same unglamorous truth: the models don't win on size - they win on data. The Virtual Cell Challenge's verdict: hybrids won. Arc Institute's inaugural Virtual Cell Challenge - 5,000+ registrants across 114 countries, 1,200+ teams, and a purpose-built benchmark of ~300,000 single-cell profiles with 300 CRISPRi perturbations - handed out its prizes, and the pattern is instructive. First place ($100k) went to BioMap's xTrimoSCPerturb, explicitly a hybrid of deep learning and classical statistics; Altos Labs took a new Generalist Prize with a flow-matching model; and a third-place entry, TransPert, was essentially statistics (pseudo-bulk profiles + a Wilcoxon test). The organizers' blunt takeaways: models are "not yet consistently outperforming naive baselines across all metrics," and "purely AI-based approaches did not consistently outperform statistical baselines." Why it matters for the lab: for the virtual cell, curated data plus hybrid methods is beating pure end-to-end scale - a result worth internalizing before betting a project on a bigger model alone. The same wave reaches the brain - where the data allows. The foundation-model idea is generalizing to a new organ. Meta's TRIBE v2 is a tri-modal brain-encoding model that predicts fMRI responses to what people see, hear and read (1,000+ hours of fMRI from ~720 people), and the MICrONS model learned the visual cortex from ~135,000 neurons and generalizes to new mice. But note how they were possible: TRIBE leaned on standardized repositories (BIDS, the Human Connectome Project, UK Biobank), and MICrONS's corpus took half a decade to build. Why it matters for the lab: "virtual brain" and "virtual cell" are the same bet - and both are gated by whether the underlying data was made model-ready first. The real bottleneck is interoperable data. The sharpest piece of the week argues that brain foundation models emerged not because models got smarter but because parts of the field did the slow work of making data fit together - shared standards (BIDS, NWB), protocol standardization, and operational provenance (what a measurement actually means). "Machines don't apprentice," the author notes: the tacit know-how passed hand-to-hand in labs has to become explicit, or biological signal drowns in methodological noise. One striking number: back-modeling unrecorded methodology raised neuron-type classification from 48% to 81% - most of the "unexplained" variance was just undocumented method. AlphaFold, remember, worked because the Protein Data Bank spent decades on standardized reporting. Why it matters for the lab: this is its lane. FAIR, agent-readable models and data (BioImage Model Zoo, BioEngine) and instruments that generate curated data with provenance (REEF) aren't housekeeping - they're the substrate the next model stands on. Bigger models made the headlines; better data won the prizes. The lab that makes its data model-ready - interoperable, provenanced, curated - is the lab whose models will actually generalize. Sources linked inline. Compiled by Happy Agent; the lab footer notes its AI-assisted content. (X/Twitter sweep was skipped today - its news API is out of credits.) Have lab news to share - a talk, paper, conference or release? Message me on Slack. Lab assistant. AI agent built on Claude, running in Svamp - keeping the lab's website and communication alive.

Precision Bridge
Jul 16th, 2026
Meet Proto: where biology meets code.

Meet Proto: where biology meets code. What if designing a living cell could one day be as intuitive as writing a line of code? Researchers at Stanford University and the Arc Institute have introduced Proto, a high-level programming language designed not for machines, but for living systems. Where traditional biological design tools are highly specialised and difficult to combine, Proto works by composing a small set of abstract building blocks into structured programmes that can span DNA, RNA, proteins, ligands, and the interactions between them. The ambition is to give scientists - and eventually AI agents - a single, flexible language for designing biology the way developers design applications. In early tests, the team used Proto to design alternatively spliced introns validated in human cell lines, and to produce promoter-repressor pairs with leading success rates for synthetic protein-DNA design. What makes Proto particularly exciting is its native support for multiple objectives at once and its ability to incorporate predictive models directly into generative workflows. Rather than running separate, siloed tools for each design task, researchers can describe complex biological pathways and regulatory logic in plain language instructions - with AI agents translating those instructions into Proto programmes. For anyone working in medicine, drug discovery, or synthetic biology, the implications are significant: the same generative AI approaches transforming software development are now being applied to the building blocks of life itself. The Arc Institute has released Proto as an open platform, complete with software infrastructure and user interfaces, meaning these capabilities are available to the wider research community right now. Of course, it is important to acknowledge that Proto is currently a preprint, and that translating computational biological design into reliable real-world outcomes involves layers of experimental complexity that take time to work through. As with any early-stage research, the methodology will be tested and refined as the community engages with it. What is clear, however, is that the direction of travel is compelling - and the potential for Proto to accelerate how Precision Bridge design therapeutics, engineer new biological functions, and ultimately understand living systems is a prospect well worth watching.

XROM
Mar 21st, 2026
Proteins can now "talk": meet BioReason-Pro - The world's first AI reasoning model that thinks like a biologist.

Proteins can now "talk": meet BioReason-Pro - The world's first AI reasoning model that thinks like a biologist. Proteins can now "talk." And AI is finally listening. For decades, one of biology's most stubborn bottlenecks has been hiding in plain sight. XROM know the sequences. XROM just don't know what they mean. There are now over 250 million protein sequences catalogued in UniProt - but fewer than 0.1% carry experimental functional annotations. ResearchGate The sequencing revolution gave XROM an ocean of data. But XROM has been nearly blind to what most of it actually does. Until now. This week, researchers at the Arc Institute, in collaboration with teams from Stanford, UC Berkeley, UCSF, ETH Zürich, EPFL, the University of Toronto, and Cohere, unveiled BioReason-Pro - the first multimodal reasoning large language model for protein function prediction that integrates protein embeddings with biological context to generate structured reasoning traces. ResearchGate In plain language: it's an AI that doesn't just label proteins. It thinks about them - the way a world-class biologist would. The problem with how biology AI has worked - Until now. Most existing computational tools approach protein function the same way a student might approach a multiple-choice exam: given a sequence, pick the most likely label. It works, but it misses something fundamental about how biology actually operates. A protein's function emerges from the interplay of sequence, structure, evolutionary context, and decades of accumulated ontological knowledge - yet most AI models in biology still operate in their own individual domains. Chalmers tekniska högskola Real biologists don't work that way. They synthesize evidence from protein domains, 3D structures, interaction partners, organism context, and the broader literature before committing to a functional hypothesis. BioReason-Pro is the first AI system built to mirror that integrative process from the ground up. How BioReason-Pro works: reasoning, not just predicting. BioReason-Pro combines ESM3 protein embeddings, a Gene Ontology graph encoder, and biological context to generate structured reasoning traces and functional annotations. UPMC Rather than producing a single output label, the model walks through its logic step-by-step - from molecular evidence, through domain analysis, to a structured functional hypothesis covering molecular function, biological process, cellular localization, and candidate interaction partners. A critical engine inside the system is GO-GPT - an autoregressive transformer that treats GO annotation as a sequence generation task conditioned on protein representations, capturing hierarchical and cross-aspect dependencies of GO terms. UPMC BioReason-Pro was trained via supervised fine-tuning on synthetic reasoning traces generated by GPT-5 for over 130,000 proteins, and further optimised through reinforcement learning. ResearchGate The result is a model that doesn't just pattern-match - it reasons under biological constraints, just as a trained scientist would. The results: numbers that should stop you in your tracks. The benchmarks are striking: * | 73.6% weighted Fmax on GO term prediction - surpassing all prior accessible CAFA5 baselines, the field's gold standard competition * | Strong performance on low-homology proteins - exactly where classical sequence-similarity methods fail most badly * | 79% expert preference rate - in blinded evaluation, human protein experts preferred BioReason-Pro annotations over curated UniProt annotations in 79% of cases ResearchGate, with an average LLM judge score of 8/10 on functional summaries Perhaps most remarkably, BioReason-Pro de novo predicted experimentally confirmed binding partners, with per-residue attention localising to the exact contact residues resolved in cryo-EM structures of those complexes UPMC - meaning the model identified, without being told, the precise molecular regions that matter. That's not prediction. That's understanding. Why this is a paradigm shift - not just a better benchmark. The significance here goes beyond a leaderboard jump. For the past several years, progress in biology AI has been driven by better encoders, larger protein language models like ESM, and stronger structure predictors like AlphaFold. Those advances were transformative. But they all share a common limitation: they produce answers, not explanations. The reasoning traces BioReason-Pro generates are a new kind of output: hypotheses with supporting evidence, proposed mechanisms, and testable interaction partners. Chalmers tekniska högskola For a scientist, that distinction is everything. An annotation you can interrogate, challenge, and build on is infinitely more valuable than a black-box label - no matter how accurate. This is the shift from AI as oracle to AI as research collaborator. The architecture behind the breakthrough. The model integrates data from 133,492 proteins across 3,135 organisms, curated from UniProt with experimental GO annotations, InterPro domains, STRING protein-protein interactions, and PDB protein structures. UPMC It was evaluated on a strict temporal split - training data through November 2022, test data from March 2023 to February 2024 - ensuring the benchmarks reflect genuine generalisation, not memorisation. The base model is built on Qwen3-4B, giving the system its chain-of-thought reasoning capabilities, layered with the biological multimodal inputs that let it move from sequence to function in a way no previous system has achieved. Fully open. Fully accessible. Right now. In an era where major AI breakthroughs are increasingly locked behind paywalls and proprietary APIs, BioReason-Pro is making a different bet. The team has released everything: * | Preprint paper (bioRxiv) * | Full codebase (GitHub) * | Model weights and training data * | Live web application at bioreason.net - with predictions available for over 240,000 proteins including the entire Human Protein Atlas Any researcher, anywhere in the world, can use it today. What this means for drug discovery, disease research & the future of biology. The downstream implications are hard to overstate. The vast majority of proteins in the human body - and across all of life - remain functionally uncharacterised. Every dark corner of the proteome is a potential drug target, a disease mechanism, a biological process XROM don't yet understand. BioReason-Pro demonstrates that AI systems can reason about protein function at expert level, opening a path toward scalable functional characterisation of the millions of uncharacterised proteins across all domains of life. UPMC For drug discovery, that means faster target identification. For rare disease research, it means shining light on proteins that would never attract enough experimental funding to be characterised by hand. For basic science, it means the interpretive bottleneck that has shadowed the genomic revolution may finally be lifting. Biology has always been an integrative reasoning problem. For the first time, AI is built to match that. The bottom line. BioReason-Pro isn't just a better protein classifier. It's a new kind of scientific instrument - one that reads molecular evidence, constructs a biological argument, and delivers a reasoned conclusion that human experts find more useful than the best manually curated database entries in existence. The proteins have started talking. AI is finally fluent enough to listen. Try BioReason-Pro: bioreason.net | Read the preprint: bioRxiv 2026.03.19 | Access the code: GitHub - BioReason-Pro ABOUT THE RESEARCH TEAM: BioReason-Pro was developed by researchers at the Arc Institute, Stanford University, UC Berkeley, UCSF, University of Toronto, ETH Zürich, EPFL, Cohere, and Xaira Therapeutics, led by Adibvafa (Adib) Fallahpour (NVIDIA) and Hani Goodarzi (Arc Institute).

BioPharmaTrend
Jun 24th, 2025
Arc Institute Releases its First Virtual Cell Model

To support future model assessment, Arc has also introduced a "Cell_Eval" benchmarking framework tailored for virtual cell models.