Full-Time

Senior HPC Storage Engineer

Updated on 9/3/2026

Deadline 9/25/26
UCLA

UCLA

Public university in Los Angeles, CA

No salary listed

Los Angeles, CA, USA

Hybrid

Remote candidates must accommodate Pacific Standard Time and may need occasional on-site presence for operational, maintenance, or project requirements.

Bachelor's, Master's, PhD

Category
DevOps & Infrastructure (1)
Required Skills
Bash
Python
High Performance Computing (HPC)
Data Science
LDAP
Git
Machine Learning
Computer Networking
SAML
Ansible
Linux/Unix

Get referred to UCLA

See people who can refer or advise you

Requirements
  • At least 7 years of experience in research, enterprise, or hyperscale storage environments with responsibility for large-scale production storage services.
  • Advanced knowledge of high-performance computing, data science, and cyberinfrastructure environments supporting large-scale research workloads.
  • Knowledge of scale-out, parallel, distributed, object, and federated storage architectures, including performance characteristics, consistency models, caching strategies, and operational tradeoffs.
  • Hands-on experience architecting, deploying, operating, or substantially supporting large-scale storage platforms such as Lustre, VAST Data, GPFS/Spectrum Scale, Ceph, BeeGFS, WekaFS, or MinIO.
  • Experience operating large-scale production storage systems, including monitoring, capacity planning, lifecycle management, upgrades, performance tuning, incident response, backup, replication, and disaster recovery.
  • Advanced Linux systems administration skills, including kernel-level troubleshooting, performance profiling, storage hardware diagnostics, and tuning across InfiniBand, RoCE, and Ethernet fabrics.
  • Experience diagnosing storage performance problems across Linux clients, metadata services, storage servers, object services, network fabrics, protocol layers, and application input/output patterns.
  • Proficiency in scripting and automation using Bash, Python, or similar languages, with familiarity with Ansible and Git.
  • Experience integrating storage systems with local, campus-wide, and federated identity providers such as Active Directory, LDAP, CILogon, OpenID Connect, Security Assertion Markup Language, or Globus Auth.
  • Ability to design and execute storage benchmarks, validate vendor claims, characterize representative scientific workloads, and document results for engineering and executive audiences.
  • Ability to communicate complex technical information clearly to technical staff, researchers, leadership, vendors, and external research and education audiences.
  • Ability to work independently and collaboratively, manage competing priorities, lead technical working groups, and sustain production service quality while delivering multi-month projects.
  • Bachelor's degree in Computer Science, Computational Science, Data Science, Engineering, or a related field, or an equivalent combination of education and experience.
  • Ability to participate in an on-call rotation for emergency hardware and software support and work occasional evenings and weekends for maintenance windows or incidents.
Responsibilities
  • Lead the design, deployment, and operation of large-scale storage systems supporting data-intensive research, artificial intelligence and machine learning, interactive computing, and multi-institutional collaborations.
  • Evaluate emerging storage technologies and lead deployment of petabyte-scale storage platforms.
  • Optimize storage performance and develop secure, reliable storage services spanning campus, cloud, and federated high-performance computing environments.
  • Help build national-scale research data infrastructure enabling seamless access to data across geographically distributed computing resources.
  • Operate and support production storage services, including monitoring, capacity planning, lifecycle management, upgrades, performance tuning, incident response, backup, replication, and disaster recovery.
  • Design and execute storage benchmarks, characterize scientific workloads, and document results for engineering and executive audiences.
  • Lead technical working groups and sustain production service quality while delivering multi-month projects.
  • Provide emergency hardware and software support through an on-call rotation and support scheduled maintenance windows or incidents.
Desired Qualifications
  • Experience with petabyte-scale storage, federated storage, multi-site research infrastructure, or high-performance computing/cloud-integrated storage services.
  • Experience with multiple storage platforms and petabyte-scale deployments.
  • Ability to design, operate, or support shared research storage services, including allocation models, user-facing service delivery, large-scale data movement, and escalation support for complex research workflows.
  • Master's degree in Computer Science, Computational Science, Data Science, Engineering, or a related field, or equivalent advanced professional experience.
  • PhD in Computer Science, Computational Science, Data Science, Engineering, or a related field, or equivalent advanced professional experience.

UCLA is the University of California, Los Angeles, a public research university in Los Angeles, California. It provides undergraduate, graduate, and professional education alongside research and public service.

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A

Get referred to UCLA

See people who can refer or advise you