Full-Time

Data Infrastructure Engineer

Glyphic Biotechnologies

Glyphic Biotechnologies

11-50 employees

Next-generation protein sequencing platform provider

No salary listed

Berkeley, CA, USA

Hybrid

Hybrid role; ~20% on-site in Berkeley, CA with flexibility for additional on-site collaboration.

Master's, PhD

Category
DevOps & Infrastructure (2)
,
Required Skills
Streamlit
Bash
Data Lake
Python
BigQuery
SQL
Machine Learning
Postgres
ETL
Infrastructure as Code (IaC)
Docker
AWS
JIRA
Confluence
Linux/Unix
Looker
Snowflake
Excel/Numbers/Sheets

Get referred to Glyphic Biotechnologies

See people who can refer or advise you

Requirements
  • MS or PhD in Computer Science, Bioinformatics, Computational Biology, Data Engineering, or a related field.
  • 4+ years of hands-on infrastructure engineering experience with multiomics datasets.
  • Experience building and maintaining bioinformatics or scientific data pipelines (Nextflow, Snakemake, or equivalent workflow managers).
  • Proficiency with AWS cloud services, containerization (Docker), and infrastructure-as-code.
  • Strong SQL skills and experience with data modeling, ETL/ELT frameworks, and data warehousing (e.g., PostgreSQL, DuckDB, BigQuery, or Snowflake).
  • Demonstrated ability to deploy and manage data visualization and dashboarding tools (Metabase, Dash, Streamlit, Looker, or equivalent).
  • Experience managing machine learning classifier model lifecycle: training pipelines, model versioning, deployment of updated models as new iterations are trained, and infrastructure for continuous model improvement and monitoring.
  • Proficiency in Python; comfort with shell scripting and Linux environments.
Responsibilities
  • Own and extend end-to-end Nextflow pipelines on AWS (Seqera Platform) that process nanopore sequencing output: basecalling (Dorado), amino acid calling, signal alignment, and ML-based amino acid classification.
  • Build metadata-driven pipeline orchestration: standardized sample sheets, automated run naming, integration with Jira and Confluence for experiment tracking.
  • Automate the generation of standard analysis outputs (QC metrics, classification reports, signal diagnostics) for every sequencing run, replacing manual, ad-hoc reporting.
  • Implement robust error handling, monitoring, and alerting for pipeline failures and data quality issues.
  • Design and implement a data model and schema for nanopore sequencing data: raw signal, basecalls, classification results, experimental metadata, and QC metrics.
  • Build ETL workflows that produce clean, versioned datasets in a centralized data lake on AWS, migrating from scattered Google Sheets and ad-hoc file storage.
  • Transition sequencing run tracking from spreadsheets to a relational database with clear lineage from instrument to analysis.
  • Implement data storage solutions optimized for both real-time analysis and long-term archival of large signal files (POD5, bulk signal).
  • Deploy and maintain data visualization tools (dashboards, interactive browsers) that allow scientists to independently explore sequencing metrics: yields, classification accuracy, plate-level comparisons, signal quality trends.
  • Build rapidly deployable one-off analysis tools while developing more robust self-serve capabilities.
  • Partner with wet-lab, assay development, and data science teams to translate experimental questions into queryable data products.
  • Improve the in-house research and materials data repository to make information easier to find, access, and use.
  • Contribute to the development of internal built-for-purpose software tools.
  • Leverage AI coding tools (Claude Code, Copilot, etc.) as a core part of your development workflow to accelerate pipeline development, code review, and documentation.
  • Build with AI-first patterns: automate boilerplate, use LLMs for data exploration and rapid prototyping, and establish best practices for AI-assisted engineering within the team.
  • Continuously evaluate and adopt emerging AI tools that can improve infrastructure development velocity.
Desired Qualifications
  • Experience with nanopore or next-generation sequencing data formats (POD5, FAST5, BAM) and analysis tools (Dorado, minimap2, samtools).
  • Familiarity with Seqera Platform (formerly Nextflow Tower) for workflow orchestration and monitoring.
  • Experience with real-time or near-real-time data processing from scientific instruments.
  • Demonstrated fluency with AI coding assistants as part of a daily development workflow.
  • Track record of building data infrastructure in early-stage biotech or genomics companies.
Glyphic Biotechnologies

Glyphic Biotechnologies

View

Glyphic Biotechnologies is advancing proteomics by building a next-generation protein sequencing platform. The core product is designed to reveal detailed information about proteins, helping researchers and medical institutions gain new insights into biology and disease. The platform supports sequencing services for researchers and may be licensed to other organizations, generating revenue from service fees and licensing agreements. Unlike many competitors that focus on smaller-scale tools, Glyphic emphasizes a first-of-its-kind sequencing capability for proteins, along with diverse deployment options (serving research labs directly and via licensing). The company's goal is to enable deeper understanding of bodily functions through protein data and to bring its sequencing technology to broader use in the life sciences through services and licensing deals.

Company Size

11-50

Company Stage

Early VC

Total Funding

$43M

Headquarters

New York City, New York

Founded

2021

Get referred to Glyphic Biotechnologies

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Glyphic raised $38 million by September 2025, extending runway for platform development.
  • Glyphic posted multiple 2026 hires in Berkeley, signaling active buildout and execution.
  • NIH and Bakar Labs visibility plus 2025 patent publications validate technical momentum.

What critics are saying

  • Glyphic still sells no commercial platform, so revenue depends on unproven hardware delivery in 2026.
  • Protein sequencing must beat mass spectrometry, or customers ignore Glyphic’s expensive workflow.
  • If ProSE misses accuracy or throughput targets, the 2025 Series A becomes a financing trap.

What makes Glyphic Biotechnologies unique

  • Glyphic’s ProSE sequences proteins de novo, one molecule at a time, without databases.
  • Glyphic claims single-molecule discrimination of all 20 amino acids and PTMs in Berkeley, 2026.
  • Glyphic holds issued patents and a 2025 nanopore peptide-sequencing filing around core chemistry.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Company Match

Unlimited Paid Time Off

Commuter Benefits

Paid Maternity and Paternity Leave

Home Office Stipend

Wellness Program

Company Equity

Employee Stock Purchase Plan

Growth & Insights and Company News

Headcount

6 month growth

6%

1 year growth

0%

2 year growth

3%
Labs of Latvia
Dec 3rd, 2024
Latvian LongeVC invests in US startup Glyphic Biotechnologies

Founded in New York in 2021 by Daniel Estandian and Joshua Yang, Glyphic Biotechnologies is developing a protein sequencing platform that uses single-molecule technology to analyze the sequence of each protein in a complex protein mixture.

BioPharmaTrend
Nov 25th, 2024
LongeVC Invests in Glyphic Biotechnologies

LongeVC has invested in Glyphic Biotechnologies, which is developing single-molecule protein sequencing technology. This innovation surpasses mass spectrometry by enabling precise sequencing of individual proteins in complex mixtures, enhancing protein analysis and broadening applications in drug discovery and diagnostics. The proteomics market, valued at over $30 billion, is set for growth with such advancements. This aligns with LongeVC's mission to support transformative healthcare technologies.

U.S. Securities and Exchange Commission
Feb 21st, 2024
SEC FORM D

The Securities and Exchange Commission has not necessarily reviewed the information in this filing and has not determined if it is accurate and complete.The reader should not assume that the information is accurate and complete.

Johnson & Johnson Innovation
Apr 5th, 2023
Awardees of the BLUE KNIGHT™ Resident QuickFire Challenge Strive Toward Inflection Points | Johnson & Johnson Innovation

By: Rachel Rath, Director, BARDA Alliance, Johnson & Johnson Innovation – JLABS & Ashim Subedee, Director, DRIVe Catalyst Office, BARDA

Bakar Labs
Jan 31st, 2023
Glyphic Biotechnologies Wins $409K SBIR Grant From NIH - Bakar Labs

Congratulations to Glyphic Biotechnologies! NIH recently reported the company had been awarded a $409 SBIR grant for their project "Single-molecule protein sequencing by iterative isolation and identification of N-terminal amino acids." The company is developing and commercializing technology to sequence proteins — instead of DNA. They're hiring!