Full-Time

AI and HPC Systems Performance Engineer

Hewlett Packard Enterprise

Hewlett Packard Enterprise

10,001+ employees

Sells enterprise hardware, software, and services

No salary listed

Bengaluru, Karnataka, India

Hybrid

Two days per week in an HPE office are required.

Master's, PhD

Category
AI & Machine Learning (1)
Required Skills
LLM
Graphics Processing Unit (GPU)
Bash
Kubernetes
Python
High Performance Computing (HPC)
CUDA
PyTorch
Machine Learning
Computer Networking
Docker
RAG
Go
Redis
Observability
C/C++
Linux/Unix
Data Analysis

Get referred to Hewlett Packard Enterprise

See people who can refer or advise you

Requirements
  • Typically 8+ years of experience.
  • Strong experience with Linux system administration and command-line environments across multiple enterprise Linux distributions.
  • Experience with modern artificial intelligence and machine learning frameworks and ecosystems including PyTorch, JAX, and Hugging Face Transformers.
  • Experience with AI model training, inference, benchmarking, performance characterization, and optimization.
  • Experience with data analysis, statistical methods, experiment design, and performance modeling techniques.
  • Experience conducting technical research and evaluating emerging AI technologies, frameworks, and hardware platforms.
  • Experience with high-performance networking technologies including InfiniBand, Remote Direct Memory Access, RDMA over Converged Ethernet, and Mellanox/NVIDIA networking solutions.
  • Strong analytical, troubleshooting, and root-cause analysis skills.
  • Proficiency in one or more programming or scripting languages such as Python, Bash, Go, or C++.
  • Experience working with complex, distributed, multi-layer software systems and AI infrastructure stacks.
  • Experience using performance profiling, tracing, observability, and benchmarking tools to analyze system and application performance.
  • Experience with containerized and orchestrated environments including Docker and Kubernetes.
  • Experience analyzing and optimizing AI workloads running on GPU-accelerated systems.
  • Experience with distributed training and inference frameworks and large-scale AI and large language model workloads.
  • Experience with GPU accelerator technologies, memory hierarchies, and AI software stacks including CUDA and NCCL.
  • Experience with distributed GPU environments and multi-node AI clusters.
  • Experience with large language model serving frameworks such as vLLM, TensorRT-LLM, or SGLang.
  • Mastery in English.
  • Experience tuning system performance in a benchmarking environment.
Responsibilities
  • Install, configure, and optimize complex AI infrastructure components including GPU servers, storage systems, high-speed networking, and AI software stacks.
  • Develop automation scripts, deployment frameworks, and Infrastructure-as-Code solutions to streamline AI platform provisioning and workload execution.
  • Perform system-level performance characterization and optimization of AI training and inference workloads on HPE platforms using GPU accelerators and distributed computing technologies.
  • Design, execute, and analyze performance benchmarks for AI and machine learning workloads, including large language models, multimodal models, Retrieval-Augmented Generation pipelines, and distributed training environments.
  • Characterize and optimize performance across multi-GPU and distributed AI environments using InfiniBand, Ethernet fabric, GPUDirect, and RDMA.
  • Capture, analyze, and interpret system telemetry, performance metrics, logs, traces, and profiling data to identify bottlenecks and optimization opportunities.
  • Develop tools, software, and automation frameworks to improve AI workload observability, performance analysis, scalability testing, and benchmark execution.
  • Collaborate with customers, partners, and internal engineering organizations to characterize, troubleshoot, and optimize AI solutions deployed on HPE infrastructure.
  • Work closely with independent software vendor, independent hardware vendor, GPU vendor, and open-source ecosystem partners to evaluate, optimize, and validate AI software and hardware solutions.
  • Evaluate emerging AI frameworks, models, accelerators, and infrastructure technologies and provide technical recommendations and performance guidance.
  • Author technical reports, white papers, reference architectures, benchmark studies, and best-practice guidance for AI performance optimization and solution design.
  • Document findings, performance issues, and optimization recommendations, and communicate technical results to engineering teams, customers, and management.
  • Provide technical leadership, mentoring, and guidance to junior engineers and contribute to performance engineering best practices.
  • Communicate project status, technical risks, and performance findings to management and stakeholders in a timely manner.
Desired Qualifications
  • 5+ years of experience in AI/ML infrastructure, performance engineering, high-performance computing, or related technical fields.
  • Experience with large-scale AI training and inference environments supporting foundation models and large language models.
  • Experience with HPE platforms, AI Factory architectures, or enterprise AI infrastructure solutions.
  • Experience with parallel and distributed storage technologies, including Weka, Lustre, BeeGFS, and GPFS.
  • Experience with AI benchmarking methodologies and industry benchmarks such as MLPerf.
  • Experience developing reference architectures, technical papers, benchmark studies, or performance guidance documentation.
  • Experience working directly with customers, partners, and cross-functional engineering teams in highly collaborative environments.
  • Experience optimizing performance across multi-GPU and multi-node AI environments using InfiniBand, RDMA, GPUDirect, or equivalent technologies.
  • Ability to work independently in globally distributed teams with minimal supervision.
Hewlett Packard Enterprise

Hewlett Packard Enterprise

View

HPE delivers enterprise IT solutions across cloud, AI, and edge computing for large organizations. It combines hardware, software, and services, with consumption-based options via HPE GreenLake and container management with HPE Ezmeral, plus Aruba networking. It differs by offering an integrated on-premises and edge-enabled stack with flexible pay-as-you-go models and active open-source engagement. Its goal is to help customers accelerate digital transformation with scalable, secure IT infrastructure across data centers, cloud, and edge.

Company Size

10,001+

Company Stage

IPO

Headquarters

Houston, Texas

Founded

1939

Get referred to Hewlett Packard Enterprise

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Fiscal Q3 2026 revenue hit $12.2 billion, up 34%, with record backlog and margins.
  • Networking revenue jumped 75% in fiscal Q3 2026, driven by routers, switching, and AI demand.
  • HPE raised FY2026 guidance after Q3, signaling momentum through 2027 enterprise AI infrastructure spending.

What critics are saying

  • August 2026 court approval forced Instant On divestiture and Mist AI licensing, weakening Juniper synergies.
  • Integration fallout and 2025-2027 restructuring cut 2,500 jobs, disrupting sales and engineering execution.
  • If Oracle or hyperscaler orders slow, HPE’s networking-led valuation collapses before Juniper integration pays off.

What makes Hewlett Packard Enterprise unique

  • HPE’s July 2025 Juniper acquisition created a broader AI networking stack than Cisco rivals.
  • GreenLake and Alletra tie storage, cloud, and operations into one enterprise procurement relationship.
  • Oracle’s September 2026 collaboration gives HPE rare exposure to giga-scale AI infrastructure buildouts.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Flexible Work Hours

Hybrid Work Options

Professional Development Budget

Wellness Program

Growth & Insights and Company News

Headcount

6 month growth

10%

1 year growth

10%

2 year growth

10%
Yahoo Finance
Sep 9th, 2026
Dell raises AI server outlook to $74B as HPE posts record results

Dell Technologies and Hewlett Packard Enterprise both reported record quarterly results driven by surging AI server demand. Dell's fiscal Q2 revenue jumped 58% year-over-year to $46.97 billion, whilst adjusted earnings per share soared 203% to $7.04. AI-optimised server revenue doubled to $16.4 billion. The company raised its fiscal 2027 revenue guidance to $192 billion and now expects adjusted EPS of $25.50. HPE's fiscal Q3 revenue rose 34% to a record $12.21 billion. Adjusted EPS climbed to $1.11 from $0.44 a year earlier. Server revenue increased 35% to $6.8 billion, whilst networking revenue surged 75% to $2.9 billion. HPE raised its full-year revenue growth forecast to 34%-37% and adjusted EPS guidance to $3.75-$3.85.

Yahoo Finance
Sep 8th, 2026
HPE grants Oracle 4M shares at $0.01 each to lock in AI network revenue

Hewlett Packard Enterprise granted Oracle warrants to purchase over 4 million shares at one penny each, linking Oracle's data centre buildout to HPE's networking revenue. The arrangement functions as a capital expenditure subsidy, binding Oracle's infrastructure spending to HPE's equity valuation. HPE's stock fell roughly 5% following cautious supply chain commentary during its earnings call, despite networking revenue rising 75% and routing revenue jumping 270% year-over-year. The warrant block is valued near $200 million. The structure incentivises Oracle to direct volume through HPE's Juniper pipeline rather than alternative providers, effectively converting a major customer into a vested stakeholder. Heavy institutional ownership in both companies, including California State Teachers Retirement System and UBS AM, reinforces the strategic partnership.

The Register
Sep 8th, 2026
HPE Alletra Storage MP B10000 R6 unifies block and file workloads with independent scaling

HPE's Alletra Storage MP B10000 Release 6, announced in May, is now generally available. The platform combines block and file storage on a single disaggregated scale-out architecture, allowing independent scaling of performance and capacity with native ransomware detection across both workload types. The B10000 uses a "shared-everything" architecture that separates compute from capacity, eliminating the need to purchase fixed controller-and-media increments. Release 6 extends this model across block and adjacent file workloads whilst maintaining a common operating environment and management plane. The platform includes AI-driven operations through HPE Data Services Cloud Console. Agentic Support Automation continuously analyses operational behaviour to detect anomalies and help drive remediation before issues escalate. HPE was recently named a Leader in the 2026 Gartner Magic Quadrant for Enterprise Storage Platforms.

Yahoo Finance
Sep 7th, 2026
Dell margins surge to 15% while HPE warns of AI squeeze despite record revenue

Dell and Hewlett Packard Enterprise both reported record revenue and raised guidance, yet investors rewarded Dell whilst punishing HPE. The divergence came down to margins under rising memory costs. HPE posted a 34% revenue increase and record 40% gross margin, but management warned margins would moderate as AI systems expand and memory shortages persist through 2027. The stock fell after hours. Dell raised full-year revenue guidance to $192 billion and demonstrated expanding margins, with its server division's operating margin jumping from 8.8% to 15% despite climbing memory prices. Management sharply raised EPS guidance, and shares surged. Following the moves, Dell now trades at a forward P/E of 20.17x versus HPE's 23.15x. Dell's earnings are expected to jump 151% in fiscal 2027.

Yahoo Finance
Sep 3rd, 2026
HPE CEO: AI demand 'exceptional,' but supply chain can't keep up

Hewlett Packard Enterprise CEO Antonio Neri told Yahoo Finance that AI demand remains "exceptional" but supply chain constraints are limiting revenue growth. HPE's networking business saw orders grow three and a half times faster than revenue, whilst traditional server orders increased 75% year over year but only delivered 35% revenue growth. Neri echoed Nvidia CEO Jensen Huang's recent comments about supply constraints hampering stronger results. The bottlenecks stem from wafer capacity and clean room yields. HPE expects some improvement in clean room operations, but Neri said the supply issues will persist until wafer capacity increases to meet demand. The company anticipates exceptional demand continuing through 2027 and beyond, driven by infrastructure build-out requiring 270 gigawatts between now and 2030.