Full-Time

High Performance Compute Systems Site Lead

LANL

Updated on 9/4/2026

Hewlett Packard Enterprise

Hewlett Packard Enterprise

10,001+ employees

Sells enterprise hardware, software, and services

Compensation Overview

$105.5k - $243k/yr

+ Variable incentives

No H1B Sponsorship

Los Alamos, NM, USA

In Person

Daily onsite work is required in Los Alamos, with additional onsite support for maintenance, incidents, and on-call responsibilities.

US Citizenship, US Top Secret Clearance Required

Bachelor's, Associate's

Category
Field Service (1)
Required Skills
Bash
SharePoint
Python
High Performance Computing (HPC)
Git
Linux/Unix
Microsoft Outlook

Get referred to Hewlett Packard Enterprise

See people who can refer or advise you

Requirements
  • US citizenship and the ability to obtain and maintain a Department of Energy Q clearance.
  • A high school diploma or equivalent with at least 7 years of relevant technical experience, or an associate or bachelor’s degree in a technical field with at least 5 years of relevant technical experience.
  • At least 5 years of hands-on experience supporting complex electronic systems, enterprise server hardware, integrated data center infrastructure, or comparable production technology environments, including diagnosing, repairing, or maintaining enterprise server components.
  • At least 3 years of hands-on experience supporting high-performance computing systems, supercomputing environments, large-scale Linux clusters, or similarly complex Linux-based compute infrastructure.
  • At least 3 years of experience providing technical leadership, mentoring, work coordination, or task direction for a multidisciplinary technical team.
  • At least 3 years of hands-on Linux system administration, production support, or troubleshooting experience with Red Hat Enterprise Linux, SUSE Linux Enterprise Server, or a comparable enterprise Linux distribution.
  • At least 2 years of experience serving as a customer-facing technical focal point and coordinating support cases, maintenance activities, escalations, or incidents in a production environment using formal ticketing, change-control, incident-management, or service-management processes.
  • Demonstrated experience using Bash, Python, or a comparable scripting language to collect system information, parse logs, automate repeatable operational tasks, analyze system data, or support troubleshooting activities.
  • Demonstrated hands-on experience safely using common hand tools, cable and fiber tools, electrostatic-discharge protection, and documented hardware service procedures to work on rack-mounted server components.
  • Demonstrated ability to apply a structured, evidence-based troubleshooting process.
  • Demonstrated ability to communicate technical information clearly to technical staff, engineers, organizational leadership, customers, and remote support organizations, including creating technical procedures, maintenance plans, support-case updates, incident summaries, troubleshooting notes, knowledge articles, executive status summaries, or customer-facing communications.
  • Experience supporting a 24x7 production environment and willingness to participate in an on-call rotation, planned after-hours maintenance, and additional onsite support during major incidents or system outages.
  • Demonstrated ability to read and interpret technical documentation, hardware diagrams, schematics, rack elevations, cable maps, network diagrams, and hardware service procedures.
  • Working proficiency with Windows or macOS and standard browser-based, collaboration, and productivity tools, including Microsoft 365, SharePoint, Slack, Outlook, and Teams.
  • Ability to work in data center and computer-room environments, lift up to 50 pounds independently and up to 75 pounds with assistance, perform rack-level and equipment-handling activities, and follow safety, security, documentation, configuration-control, and information-protection requirements.
Responsibilities
  • Provide day-to-day technical leadership and guidance to onsite hardware engineers, Linux system administrators, and software analysts while coordinating work across inventory specialists and remote engineering resources.
  • Set and communicate daily and weekly technical priorities based on system health, support cases, maintenance, customer priorities, operational risk, service commitments, and staffing.
  • Serve as the primary onsite technical focal point for system support, technical escalations, maintenance activities, and service-delivery risks.
  • Maintain visibility into system health, support-case status, technical risks, maintenance actions, ownership, and outstanding commitments, and lead routine operational reviews.
  • Ensure support cases contain accurate technical details, diagnostic evidence, business impact, troubleshooting history, ownership, and next actions.
  • Drive timely escalation through established processes and engage next-tier support, engineering, product teams, and other resources needed for diagnosis and resolution.
  • Plan and coordinate system upgrades, maintenance windows, installations, acceptance activities, and other production-impacting work in accordance with change-control and support processes.
  • Confirm staffing, parts, tools, test equipment, documentation, communications, escalation contacts, rollback plans, and contingency coverage for planned work.
  • Lead the onsite response during major incidents by aligning priorities with the customer incident lead, organizing resources, maintaining communications, tracking actions and decisions, and escalating appropriately.
  • Coordinate root-cause analysis, corrective actions, lessons learned, and documentation updates after significant incidents or recurring issues.
  • Identify and promptly escalate risks to system availability, service-level commitments, maintenance schedules, operational readiness, or customer satisfaction.
  • Coordinate onsite parts inventory, repair materials, tools, test equipment, and other company-owned resources using approved business systems and controls.
  • Build and maintain working relationships with onsite team members, remote engineering organizations, leadership, and customer stakeholders.
  • Facilitate team coordination discussions to organize work, confirm ownership, review progress, surface blockers, and ensure commitments are completed or escalated.
  • Provide technical mentoring, coaching, and practical guidance without assuming formal people-management authority.
  • Promote accountability, disciplined troubleshooting, accurate documentation, safe work practices, knowledge sharing, and professional customer engagement.
  • Track required company, customer, security, safety, and technical training and escalate overdue or at-risk requirements.
  • Coordinate onsite coverage using approved schedules and planned leave information and escalate coverage gaps.
  • Provide fact-based observations to the delivery manager regarding technical performance, development needs, recognition opportunities, and issues requiring formal management attention.
  • Coordinate site-specific onboarding and operational readiness for new team members.
  • Maintain site procedures, contact lists, escalation paths, team schedules, operational references, and other support information.
  • Maintain a clean, safe, secure, and organized working environment.
  • Participate in the on-call rotation and provide additional onsite support for 24x7 operations, planned maintenance, system outages, and major incidents.
  • Use Linux command-line and diagnostic tools to analyze processes, filesystems, services, permissions, network state, system logs, hardware telemetry, and system health.
  • Lead and participate in monitoring, diagnosis, maintenance, and restoration of HPC compute, interconnect, storage, management, power, and cooling infrastructure.
  • Apply evidence-based troubleshooting that correlates Linux logs, hardware telemetry, network state, firmware information, and case history to isolate faults and determine corrective actions or escalation.
  • Diagnose and support repair of enterprise server hardware, compute nodes, management components, interconnect and network components, storage, power systems, cabling, and integrated HPC equipment.
  • Perform or assist with hardware replacement, cable and fiber management, rack-level work, equipment installation, electrostatic-discharge controls, and other data center activities.
  • Use out-of-band management controllers and interfaces, including BMCs, Redfish, and IPMI, to assess hardware health, validate firmware, review environmental conditions and event logs, manage power state, and verify component status.
  • Read and interpret system documentation, hardware diagrams, rack elevations, network diagrams, cable maps, schematics, and service procedures.
  • Support new-system installation, expansion, integration, acceptance testing, hardware and firmware upgrades, and transition-to-operations activities.
  • Create and maintain site documentation, troubleshooting procedures, maintenance plans, workflows, technical checklists, incident records, and knowledge articles.
  • Use Bash, Python, Git, and approved scripting, version-control, and collaboration tools to collect information, automate tasks, analyze system data, and maintain operational documentation and configuration references.
Desired Qualifications
  • Prior Department of Energy Q, Department of Defense Top Secret, or comparable federal security-clearance experience; an active Department of Energy Q clearance is strongly preferred.
  • Hands-on experience supporting HPE Cray EX systems, HPE Cray supercomputing platforms, or other leadership-class and exascale supercomputing environments.
  • Experience with enterprise Linux diagnostics and administration, including system services, package management, network troubleshooting, log correlation, performance analysis, system-health assessment, and automation.
  • Experience with HPE Slingshot interconnects, InfiniBand fabrics, high-speed Ethernet, or other large-scale HPC networking technologies.
  • Experience with high-speed network diagnostics, optical transceivers, copper and fiber cabling, link analysis, topology review, and fault isolation.
  • Experience with liquid-cooled HPC infrastructure, cooling distribution units, high-density compute cabinets, direct-liquid cooling, or related facility interfaces.
  • Experience using Redfish, IPMI, BMC interfaces, or comparable out-of-band management technologies.
  • Experience supporting NVIDIA GB200 or GB300 NVL72 systems, NVIDIA DGX platforms, high-density GPU systems, accelerated computing platforms, or other rack-scale AI infrastructure.
  • Experience supporting parallel filesystems, enterprise storage platforms, cluster-management services, workload managers or schedulers, or HPC monitoring systems.
  • Project-management or technical workstream leadership experience coordinating upgrades, installations, maintenance windows, acceptance activities, or multi-party technical projects.
  • Experience with formal incident management, problem management, change management, service-level management, or ITIL-aligned service-delivery practices.
  • Experience using Git or a comparable version-control platform for scripts, configuration files, procedures, and technical documentation.
  • Experience working at a Department of Energy national laboratory, Department of Defense facility, government research organization, or another highly regulated customer environment.
  • Relevant industry certifications such as CompTIA Linux+, Server+, Network+, Security+, Red Hat Certified System Administrator, ITIL Foundation, or equivalent technical certifications.
Hewlett Packard Enterprise

Hewlett Packard Enterprise

View

HPE delivers enterprise IT solutions across cloud, AI, and edge computing for large organizations. It combines hardware, software, and services, with consumption-based options via HPE GreenLake and container management with HPE Ezmeral, plus Aruba networking. It differs by offering an integrated on-premises and edge-enabled stack with flexible pay-as-you-go models and active open-source engagement. Its goal is to help customers accelerate digital transformation with scalable, secure IT infrastructure across data centers, cloud, and edge.

Company Size

10,001+

Company Stage

IPO

Headquarters

Houston, Texas

Founded

1939

Get referred to Hewlett Packard Enterprise

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Fiscal Q3 2026 revenue hit $12.2 billion, up 34%, with record backlog and margins.
  • Networking revenue jumped 75% in fiscal Q3 2026, driven by routers, switching, and AI demand.
  • HPE raised FY2026 guidance after Q3, signaling momentum through 2027 enterprise AI infrastructure spending.

What critics are saying

  • August 2026 court approval forced Instant On divestiture and Mist AI licensing, weakening Juniper synergies.
  • Integration fallout and 2025-2027 restructuring cut 2,500 jobs, disrupting sales and engineering execution.
  • If Oracle or hyperscaler orders slow, HPE’s networking-led valuation collapses before Juniper integration pays off.

What makes Hewlett Packard Enterprise unique

  • HPE’s July 2025 Juniper acquisition created a broader AI networking stack than Cisco rivals.
  • GreenLake and Alletra tie storage, cloud, and operations into one enterprise procurement relationship.
  • Oracle’s September 2026 collaboration gives HPE rare exposure to giga-scale AI infrastructure buildouts.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Flexible Work Hours

Hybrid Work Options

Professional Development Budget

Wellness Program

Growth & Insights and Company News

Headcount

6 month growth

10%

1 year growth

10%

2 year growth

10%
Yahoo Finance
Sep 9th, 2026
Dell raises AI server outlook to $74B as HPE posts record results

Dell Technologies and Hewlett Packard Enterprise both reported record quarterly results driven by surging AI server demand. Dell's fiscal Q2 revenue jumped 58% year-over-year to $46.97 billion, whilst adjusted earnings per share soared 203% to $7.04. AI-optimised server revenue doubled to $16.4 billion. The company raised its fiscal 2027 revenue guidance to $192 billion and now expects adjusted EPS of $25.50. HPE's fiscal Q3 revenue rose 34% to a record $12.21 billion. Adjusted EPS climbed to $1.11 from $0.44 a year earlier. Server revenue increased 35% to $6.8 billion, whilst networking revenue surged 75% to $2.9 billion. HPE raised its full-year revenue growth forecast to 34%-37% and adjusted EPS guidance to $3.75-$3.85.

Yahoo Finance
Sep 8th, 2026
HPE grants Oracle 4M shares at $0.01 each to lock in AI network revenue

Hewlett Packard Enterprise granted Oracle warrants to purchase over 4 million shares at one penny each, linking Oracle's data centre buildout to HPE's networking revenue. The arrangement functions as a capital expenditure subsidy, binding Oracle's infrastructure spending to HPE's equity valuation. HPE's stock fell roughly 5% following cautious supply chain commentary during its earnings call, despite networking revenue rising 75% and routing revenue jumping 270% year-over-year. The warrant block is valued near $200 million. The structure incentivises Oracle to direct volume through HPE's Juniper pipeline rather than alternative providers, effectively converting a major customer into a vested stakeholder. Heavy institutional ownership in both companies, including California State Teachers Retirement System and UBS AM, reinforces the strategic partnership.

The Register
Sep 8th, 2026
HPE Alletra Storage MP B10000 R6 unifies block and file workloads with independent scaling

HPE's Alletra Storage MP B10000 Release 6, announced in May, is now generally available. The platform combines block and file storage on a single disaggregated scale-out architecture, allowing independent scaling of performance and capacity with native ransomware detection across both workload types. The B10000 uses a "shared-everything" architecture that separates compute from capacity, eliminating the need to purchase fixed controller-and-media increments. Release 6 extends this model across block and adjacent file workloads whilst maintaining a common operating environment and management plane. The platform includes AI-driven operations through HPE Data Services Cloud Console. Agentic Support Automation continuously analyses operational behaviour to detect anomalies and help drive remediation before issues escalate. HPE was recently named a Leader in the 2026 Gartner Magic Quadrant for Enterprise Storage Platforms.

Yahoo Finance
Sep 7th, 2026
Dell margins surge to 15% while HPE warns of AI squeeze despite record revenue

Dell and Hewlett Packard Enterprise both reported record revenue and raised guidance, yet investors rewarded Dell whilst punishing HPE. The divergence came down to margins under rising memory costs. HPE posted a 34% revenue increase and record 40% gross margin, but management warned margins would moderate as AI systems expand and memory shortages persist through 2027. The stock fell after hours. Dell raised full-year revenue guidance to $192 billion and demonstrated expanding margins, with its server division's operating margin jumping from 8.8% to 15% despite climbing memory prices. Management sharply raised EPS guidance, and shares surged. Following the moves, Dell now trades at a forward P/E of 20.17x versus HPE's 23.15x. Dell's earnings are expected to jump 151% in fiscal 2027.

Yahoo Finance
Sep 3rd, 2026
HPE CEO: AI demand 'exceptional,' but supply chain can't keep up

Hewlett Packard Enterprise CEO Antonio Neri told Yahoo Finance that AI demand remains "exceptional" but supply chain constraints are limiting revenue growth. HPE's networking business saw orders grow three and a half times faster than revenue, whilst traditional server orders increased 75% year over year but only delivered 35% revenue growth. Neri echoed Nvidia CEO Jensen Huang's recent comments about supply constraints hampering stronger results. The bottlenecks stem from wafer capacity and clean room yields. HPE expects some improvement in clean room operations, but Neri said the supply issues will persist until wafer capacity increases to meet demand. The company anticipates exceptional demand continuing through 2027 and beyond, driven by infrastructure build-out requiring 270 gigawatts between now and 2030.