Full-Time

Lead Systems Engineer

HPC

Updated on 9/12/2026

Deadline 6/4/27
Princeton University

Princeton University

Research university in Princeton, NJ

Compensation Overview

$135k - $150k/yr

Princeton, NJ, USA

In Person

Bachelor's

Category
IT Operations (1)
Required Skills
Bash
Python
High Performance Computing (HPC)
Computer Networking
Perl
Linux/Unix

Get referred to Princeton University

See people who can refer or advise you

Requirements
  • At least 10 years of experience managing advanced research computing systems.
  • Strong expertise in Linux system administration, installation, and troubleshooting.
  • Advanced experience writing scripts in Bash, Python, and/or Perl.
  • Proficiency managing networking in high-performance computing environments.
  • Strong experience managing software in an advanced research computing environment.
  • Experience supporting scheduling and managing jobs with SLURM in large-scale computing environments.
  • Strong oral and written communication skills, with the ability to proactively engage peers and communicate effectively across a diverse stakeholder community.
  • Strong ability to solve complex system and infrastructure problems and share expertise with colleagues at all levels.
  • Demonstrated ability to collaborate across teams to solve systems and infrastructure challenges while aligning day-to-day operational needs with longer-term technical and organizational goals.
  • Ability to maintain personal, proprietary, and confidential data in the strictest confidence and follow procedures for its privacy, security, and proper use.
  • A bachelor's degree in a related field or equivalent experience.
Responsibilities
  • Design, maintain, troubleshoot, and refine advanced high-performance computing and artificial intelligence cluster infrastructure, including high-performance interconnects, cluster schedulers, and configuration management across research systems.
  • Partner with Advanced Data and Storage Management colleagues to align scratch filesystem and data-management designs with cluster designs.
  • Develop data-transfer pathways and networks to support artificial-intelligence-driven computing workloads.
  • Establish and maintain best practices for cluster management and usage to support artificial-intelligence-driven workloads.
  • Develop documentation for users and technical staff for use by the broader community.
  • Develop, enhance, and expand monitoring infrastructure and related protocols for research computing systems.
  • Plan and implement scheduled operations maintenance, including during off-hours.
  • Define and drive institutional technical strategy for advanced artificial intelligence and data-intensive high-performance computing.
  • Anticipate and solve novel and complex problems, determine project objectives and requirements, and develop standards and governance for research computing platforms.
  • Identify, evaluate, and pilot researcher-facing systems that accelerate research using artificial intelligence.
  • Lead implementation and expand adoption of modern, automation-driven infrastructure and cluster-management practices.
  • Promote institution-wide collaboration as a community expert advising and working with faculty, researchers, and vendors on emerging trends and challenges in artificial-intelligence-enabled research computing.
  • Provide technical mentorship to systems specialists and analysts by sharing designs and operational expertise across data systems and high-performance computing and artificial-intelligence infrastructure.
  • Contribute to the strategic vision for high-performance computing and artificial-intelligence systems and advise senior leadership and stakeholders on strategic investments, risks, and opportunities related to research infrastructure.
  • Monitor high-performance computing clusters, networks, and storage systems for abnormalities and resolve issues.
  • Analyze and solve problems in Linux and high-performance computing and artificial-intelligence environments involving software, data, and job submissions.
  • Use scripting and programming tools to troubleshoot issues.
  • Participate in a mandatory on-call rotation involving infrequent off-hour and weekend duty.
Desired Qualifications
  • Experience working in academic and research settings.
  • Experience supporting artificial-intelligence-driven research in open and secure computing environments.
  • Familiarity with data-transfer technologies such as Globus for transferring large datasets.
  • Experience using and supporting parallel file systems commonly used in high-performance computing and artificial-intelligence systems.
  • Experience supporting unstructured data in high-performance computing and artificial-intelligence environments.

Princeton University is an independent research university in Princeton, New Jersey. It emphasizes undergraduate and doctoral education alongside scholarship, research, and teaching.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A

Get referred to Princeton University

See people who can refer or advise you