Full-Time

Principal Platform Engineer

Observability, Cipe

Palo Alto Networks

Palo Alto Networks

10,001+ employees

Firewall and cloud security provider

Compensation Overview

$147k - $237.5k/yr

+ Restricted Stock Units + Bonus

No H1B Sponsorship

California, USA

In Person

On-site role based in California; remote work not available.

Category
DevOps & Infrastructure (1)
Required Skills
LLM
Claude
Kubernetes
Python
OpenTelemetry
Prometheus
Helm

Get referred to Palo Alto Networks

See people who can refer or advise you

Requirements
  • 7+ years of software engineering, platform engineering, infrastructure engineering, or SRE experience, with significant experience building production-grade distributed systems.
  • Deep hands-on experience with observability systems, including metrics, logs, traces, profiling, dashboards, synthetics, alerting, and incident workflows.
  • Strong expertise with OpenTelemetry, including SDKs, Collector pipelines, exporters, processors, receivers, semantic conventions, and instrumentation patterns.
  • Strong experience with Prometheus-compatible metrics, Alertmanager, scraping, cardinality management, federation, and remote write patterns.
  • Hands-on experience with distributed tracing systems such as Jaeger or similar technologies.
  • Experience with continuous profiling technologies.
  • Strong experience with synthetic monitoring and proactive availability testing, including API checks, browser-based checks, blackbox monitoring, dependency checks, and integration with alerting and SLO workflows.
  • Strong Kubernetes experience, including workload monitoring, service discovery, operators/controllers, Helm, resource management, cluster observability, and multi-tenant platform patterns.
  • Strong Python engineering skills, including building internal tools, automation, integrations, services, and instrumentation libraries.
  • Hands-on experience building real solutions, tools, and developer workflows using modern AI coding agents such as Claude, Codex, or equivalent — including prompt design, skill/tool/MCP authoring, agent orchestration, and integrating LLMs into production engineering systems.
  • Practical understanding of how to design AI-friendly platforms: structured APIs, machine-readable runbooks, telemetry schemas, and skills/tools that allow both humans and AI agents to operate observability effectively.
  • Experience designing and operating high-scale, highly available infrastructure systems.
  • Strong understanding of SLOs, SLIs, error budgets, incident response, on-call practices, production readiness, and reliability engineering principles.
  • Experience writing clear technical design documents, RFCs, standards, operational runbooks, and architecture recommendations.
  • Ability to influence teams through technical depth, collaboration, mentorship, and pragmatic decision-making.
Responsibilities
  • Design and lead the evolution of a modern observability platform using OpenTelemetry, Prometheus, Jaeger, Alertmanager, and related CNCF ecosystem tools.
  • Define architecture standards for telemetry collection, processing, storage, querying, visualization, alerting, retention, and governance.
  • Build scalable systems for metrics, distributed tracing, continuous profiling, log aggregation, synthetic monitoring, service health monitoring, and reliability analytics.
  • Establish best practices for instrumentation across services, infrastructure, Kubernetes workloads, CI/CD systems, and developer platforms.
  • Evaluate trade-offs around data cardinality, sampling, storage cost, retention, query performance, multi-tenancy, reliability, and operational complexity.
  • Make pragmatic recommendations on open source, self-managed, managed-service, and hybrid observability approaches.
  • Create paved-road observability patterns that help engineering teams instrument, monitor, debug, and operate services with minimal friction.
  • Lead adoption and standardization of OpenTelemetry across applications, services, infrastructure, and platform components.
  • Design and implement telemetry pipelines using OpenTelemetry Collector, exporters, processors, receivers, connectors, and custom extensions where needed.
  • Define conventions for traces, metrics, logs, spans, attributes, resources, service names, correlation IDs, and semantic conventions.
  • Build libraries, SDK wrappers, golden paths, and internal tooling to simplify observability instrumentation for engineering teams.
  • Architect metrics systems using Prometheus-compatible formats, PromQL, remote write, federation, scraping strategies, service discovery, recording rules, and long-term storage backends.
  • Design alerting frameworks that reduce noise, improve signal quality, and align with SLOs, SLIs, error budgets, and incident response practices.
  • Create reusable alerting patterns for Kubernetes, infrastructure, applications, APIs, databases, queues, event-driven systems, and distributed services.
  • Define standards for dashboarding, runbooks, escalation policies, alert ownership, and production readiness.
  • Partner with SRE and engineering teams to mature monitoring practices and improve service reliability.
  • Build observability capabilities for Kubernetes environments, including cluster monitoring, workload telemetry, service mesh visibility, ingress and egress monitoring, and node-level insights.
  • Develop and maintain Helm charts, Kubernetes manifests, operators, sidecars, agents, DaemonSets, and deployment automation for observability components.
  • Work with platform teams to ensure observability systems are reliable, secure, multi-tenant, highly available, and easy to operate.
  • Define standards for resource usage, scaling, upgrades, failover, backup, disaster recovery, access control, and tenant isolation for observability infrastructure.
  • Support observability across multi-cluster, multi-region, and hybrid cloud environments where applicable.
  • Design and build AI-enabled observability workflows that allow both humans and AI agents to investigate incidents, query telemetry, summarize signals, and propose remediations.
  • Define and publish reusable AI skills, agents, and tools (e.g., Claude skills, Codex tools, MCP servers, structured prompts) that encode observability best practices and make platform capabilities consumable by engineering teams and autonomous agents.
  • Build paved-road AI integrations for triage, alert summarization, root-cause analysis, log/trace exploration, runbook generation, dashboard authoring, and post-incident review.
  • Establish standards for grounding AI agents in authoritative telemetry, runbooks, and service catalogs, with strong guardrails around accuracy, safety, cost, and auditability.
  • Use AI coding tools (Claude, Codex, and equivalents) as a first-class part of the engineering workflow — for code generation, refactoring, instrumentation rollouts, migrations, and platform automation — and define patterns the broader team can adopt.
  • Partner with platform, SRE, and product teams to evolve observability from human-only dashboards toward agent-assisted, self-serve reliability operations.

Palo Alto Networks provides hardware, software, and subscription-based security solutions to protect organizations from cyber threats. Its products include network firewalls (Strata), cloud security (Prisma Cloud), and AI-powered security operations (Cortex), which work together to secure on-premises and cloud environments, manage identity and access, and detect threats. The company differentiates itself by offering an integrated, end-to-end security stack that combines hardware, software licenses, and ongoing services across enterprises, SMBs, and government clients. Its goal is to deliver comprehensive protection for networks, data, and applications as organizations move between data centers and cloud environments.

Company Size

10,001+

Company Stage

IPO

Headquarters

Santa Clara, California

Founded

2015

Get referred to Palo Alto Networks

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • Q3 FY26 revenue rose 31% to $3.0 billion, beating estimates on June 2, 2026.
  • Next-generation security ARR grew 60% to $8.1 billion, signaling durable platform demand.
  • Lumen launched Cortex XSIAM MDR, validating enterprise adoption and ecosystem pull around Palo Alto.

What critics are saying

  • China opened a cybersecurity review on August 6, 2026; procurement freezes follow quickly.
  • CyberArk and Chronosphere integration increases execution risk, especially as acquisitions drive reported growth.
  • Palo Alto still posted a $177 million GAAP loss in Q3 FY26; valuation relies on flawless scaling.

What makes Palo Alto Networks unique

  • Palo Alto’s platformization bundles firewall, SASE, XSIAM, AI security, and observability into one stack.
  • Prisma AIRS hit 300 customers by June 2026, accelerating AI-native security leadership.
  • Chronosphere pushed observability above $300 million ARR in Q3 FY26, expanding beyond cybersecurity.

Help us improve and share your feedback! Did you find this helpful?

Benefits

FLEXBenefits

Healthcare

Wellness

Development

Financial: Traditional & Roth 401(k) options

Time Off

Other Perks

Growth & Insights and Company News

Headcount

6 month growth

3%

1 year growth

-1%

2 year growth

-1%
Yahoo Finance
Aug 11th, 2026
Palo Alto Networks stock surges 142% as AI-driven cyber threats fuel firewall demand and $3.35B Chronosphere deal

Palo Alto Networks' stock has surged 142% over six months to $385.04, but the catalysts were visible in company reports well before the market reacted. The cybersecurity firm disclosed in May 2025 that its researchers simulated a complete ransomware attack in under 25 minutes using AI, signalling a shift in threat landscape. Management responded by acquiring observability platform Chronosphere for $3.35 billion in January 2026 and raising its fiscal 2030 next-generation security ARR target to $20 billion from $15 billion. The thesis materialised in June 2026 results: next-generation security ARR grew 28% year-over-year, whilst next-generation firewall bookings jumped nearly 40%. Hardware delivered its strongest quarter in a decade on fifth-generation appliances and AI data centre deployments, exceeding management's earlier mid-single-digit growth expectations for fiscal 2026.

Yahoo Finance
Aug 9th, 2026
Palo Alto Networks triples to nearly $300B market value in 12 months on AI security demand

Palo Alto Networks' market value has surged from $113 billion a year ago to nearly $295 billion, a gain of more than 150%. The cybersecurity company acquired identity-security firm CyberArk and observability company Chronosphere, whilst management reports accelerating bookings from customers securing AI deployments. Revenue grew 31% year over year to $3.0 billion in the fiscal third quarter ended 30 April, though $388 million came from acquisitions. Organic revenue growth was approximately 14%. Next-generation security annual recurring revenue reached $8.1 billion, up 60% year over year, with acquisitions contributing $1.6 billion. Excluding acquisitions, growth was 28%. Adjusted earnings per share rose 6% to $0.85, whilst under GAAP accounting the company posted a $177 million quarterly loss. Management maintains guidance for 24% revenue growth in fiscal 2026 and a 40% adjusted free cash flow margin by fiscal 2028.

Yahoo Finance
Aug 6th, 2026
China launches security review into Palo Alto Networks products, mirroring Micron's 2023 case

China has launched a cybersecurity review into Palo Alto Networks' products sold in the country, mirroring the regulatory process that led to restrictions on Micron products in 2023. The investigation aims to ensure safe operation of critical infrastructure and prevent security risks. PANW stock traded flat in early hours despite the news, holding near its record high of over $376. The shares have nearly doubled this year. The financial impact may be limited. Palo Alto does not report China revenue separately, but the Asia-Pacific region accounted for approximately 12% of its fiscal 2025 revenue, or roughly $1.1 billion out of $9.2 billion total. China likely represents only a portion of that broader figure, which includes Japan, Australia, India, and Southeast Asia.

Yahoo Finance
Aug 1st, 2026
Cramer backs Palo Alto and CrowdStrike in cybersecurity's AI era

Jim Cramer expressed confidence in cybersecurity stocks, particularly Palo Alto Networks and CrowdStrike, predicting continued strong performance. Palo Alto's shares have risen 91% over the past year, whilst CrowdStrike's stock has gained 67%. Cramer highlighted both companies' involvement with the Open Secure AI Alliance, established by NVIDIA to develop open technologies for AI development. He believes cybersecurity firms will benefit from AI's demand for data and computing resources, alongside threats to US infrastructure. CrowdStrike reported 26% annual revenue growth to $1.39 billion, with annual recurring revenue reaching $4.92 billion through a 23% increase. However, analysts debate whether the company's 151 forward price-to-earnings multiple is justified, particularly as it remains unprofitable on a GAAP basis. Palo Alto Networks trades at a 78.74 P/E multiple.

PR Newswire
Jul 21st, 2026
Palo Alto Networks to acquire Embrace, adds digital experience monitoring to $300M observability platform

Palo Alto Networks announced plans to acquire Embrace, a provider of user-focused observability, to add Real User Monitoring capabilities to its platform. The cybersecurity firm is also introducing Synthetics, a new feature for proactively validating application performance. The additions will extend Palo Alto Networks' Observability platform to Digital Experience Monitoring. This will give customers unified visibility from end-user interactions to backend infrastructure. Following its acquisition of Chronosphere in January 2026, Palo Alto Networks' Observability platform surpassed $300 million in annual recurring revenue in Q3 FY26. The company was named a leader in Gartner's Magic Quadrant for Observability Platforms for the third consecutive year. The Embrace acquisition is subject to customary closing conditions and is expected to close in Palo Alto Networks' first quarter of fiscal 2027.