
Work Here?
What PagerDuty does: PagerDuty provides an incident management platform that helps organizations detect and resolve IT issues quickly to minimize disruption. How its product works: It integrates with monitoring tools and IT systems to detect incidents in real time, then sends alerts to the right people, enforces on-call rotations, and guides incident resolution through automated workflows and escalations. How it differentiates from competitors: It focuses on on-call management and real-time incident response across many integrations, offering a scalable, subscription-based platform with configurable alerting, escalation policies, and professional services to support faster recovery. What the company's goal is: To reduce downtime and maintain the reliability and performance of digital services for organizations across industries by providing reliable, scalable incident management.
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
IPO
Headquarters
San Francisco, California
Founded
2009
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$741.7M
Above
Industry Average
Funded Over
8 Rounds
Health, AD&D, Disability, Vision, Life, and Dental Insurance
Paternity and Maternity Leave
Employee Assistance Program
PTO (Vacation / Personal Days)
Sick Time
Remote Work
Adoption Assistance
401(k)
Employee Stock Purchase Program
Flexible Spending Account
Student Loan Repayment Plan
SRE Agent Enhancements: faster triage, greater access controls, deeper system connectivity. This blog post is part of PagerDuty's ongoing series on how PagerDuty, Inc. is helping customers navigate their journey towards autonomous operations. Read on to learn about how recent SRE Agent Enhancements build towards this vision. During an incident, everything is competing for attention at once. Responders lose time swiveling between tools, insights gathered by AI stay siloed instead of feeding into the next decision, and the pressure to move fast means learnings rarely stick. The same issues creep back in a few weeks later, and the cycle starts over. Earlier this year, PagerDuty introduced SRE Agent as a virtual responder, one that gathers signals from across your stack to help teams triage, diagnose, and remediate, using memory from past incidents and continuous learning to improve future responses. Since then, PagerDuty, Inc. has been rolling out enhancements that make the agent faster to configure, easier to trust, and more capable the moment an incident fires. Here's what's new. Triage before a human even looks. SRE Agent can now be intelligently triggered through Escalation Policies (EA) or incident workflows (GA). Configure it to jump into action the moment an incident triggers, or set criteria based on priority or severity, and the agent joins the incident pre-armed with triage data and memory of past incidents. That means investigation and analysis can be well underway before a responder ever acknowledges the page. When you finally do open the incident, you're not starting from zero. A faster way to extend the agent. PagerDuty, Inc. also introduced a new configuration experience for agent connectors, tools, and skills. Connectors (GA) plug the agent into third-party data sources like Grafana, New Relic, and Datadog through MCP or API, just enter credentials and authorize. Tools (GA) let the agent retrieve logs, metrics, and traces from observability platforms like Datadog, or pull context from knowledge bases like Confluence and GitHub. Skills (EA) arm the agent with custom instructions and domain expertise tailored to your environment, and teams can create them directly from Claude or PagerDuty for use in Slack or the PagerDuty web platform. Together, these let SRE Agent deduce troubleshooting steps before a human even opens the incident. Governance built for enterprise rollout. Customers told PagerDuty, Inc. they wanted more control over which teams could use agents, and how that access scales across the org. PagerDuty Advance team-level permissions (GA) let you scope AI to specific teams, giving admins the governance layer needed to roll out agentic AI with confidence rather than guesswork. Recommendations, with the reasoning behind them. Beyond investigation and diagnosis, SRE Agent can now recommend the right course of action through Recommended Incident Workflows (GA). It analyzes your existing configured workflows and suggests the one that best fits the current incident, along with the reasoning behind the call. That reasoning matters as much as the recommendation itself. Responders don't just get told what to do, they see why the agent landed there, so they can validate the call and build trust in the system over time. Seeing it come together. Here's what that looks like end to end: an incident triggers, and SRE Agent, already assigned through the escalation policy, begins autonomous triage immediately. By the time a responder opens the incident, the agent has already gathered context, investigated likely causes, and identified a recommended workflow, complete with reasoning drawn from how similar incidents were resolved in the past. The responder reviews the recommendation, runs the workflow, and the agent confirms the fix worked and resolves the incident. It even generates a new runbook, so remediation is faster the next time a similar issue comes up. That's the flywheel behind Autonomous Operations: today's incident becomes tomorrow's prevention, powered by data at scale, intelligent automation, and a system that keeps getting smarter. Try it yourself. These capabilities are rolling out now, with several available as part of Early Access. Watch the full demo to see in action the SRE Agent enhancements for faster triage, greater access controls, and deeper system connectivity. Want early access SRE Agent on Escalation Policies? Sign up at pagerduty.com/early-access.
How PagerDuty powers intelligent incident response for the Golden State Warriors, UWM, and Arize. The best incident is the one your customers never notice. A payment clears, a stream doesn't buffer, a loan closes on time - all because upstream, your team diagnosed and resolved the problem before it spread. But that's getting harder to pull off. PagerDuty is already ahead of that curve - and customers are seeing compounding results. With 59% of organizations already actively incorporating AI into operational workflows, applications are generating more signals (and more novel failure modes) than legacy monitoring systems were ever designed to handle. That means more noise for your team to sift through to reach the root of the issue. Meanwhile, customers get frustrated, and you creep closer to the limits of your SLAs. The enterprises winning in this era aren't just collecting more signals. They're making them actionable. PagerDuty's Operations Cloud uses AI Ops - including intelligent triage, automated runbooks, and the PagerDuty SRE Agent - to accelerate alert triage, route incidents to the right responders with full context, and resolve routine issues automatically. Now, teams can detect, fix, and even prevent incidents before they impact end users. Here's what the benefits look like for three PagerDuty customers. Catch critical issues before they affect service. The Golden State Warriors run one of the most technology-forward operations in sports. At Chase Center, their home arena in the Bay Area, the fan experience spans online ticket sales, in-venue food and beverage transactions, and a website that serves more than 70,000 active users a month. Fans expect a seamless experience, so every system has to perform perfectly during the game, when there's no option to delay service. Before PagerDuty, executives or fans flagged 80% of issues before the Warriors' IT team ever knew about them. In one case, a bug buried in third-party code slipped through testing and reversed an entire category of transactions before anyone caught it the next day. With PagerDuty, their system now tracks logs across the Warriors' digital platforms, surfaces irregular patterns, and alerts the team the moment something looks off - sending them to the people who can fix them before they cascade into a game day service failure. PagerDuty also helped the team map its incident response from end to end, defining escalation paths for both critical and non-urgent incidents. Each alert now gets prioritized, so the team applies the right level of response every time. "We need to be able to respond quickly, and PagerDuty allows us to inform the right person at the right time so that we're able to intervene and get a resolution as fast as we can," says Nick Manning, Senior Director of Consumer Product and Emerging Technologies at the Golden State Warriors. Now the team catches issues before executives or fans do, maintaining a seamless fan experience instead of scrambling to react. Modernize systems for better communication. United Wholesale Mortgage (UWM) is ranked among the nation's top mortgage lenders. When their service is disrupted, borrowers aren't just frustrated - they're delayed from closing on their homes. But for years, UWM's incident response workflows actually got in the way of their mission. Before PagerDuty, brokers reported issues to the help desk, who then tracked down a few key people by email or text, then tagged the rest of the IT floor in Microsoft Teams (day or night). Each manual handoff burned critical minutes and fragmented context, leaving leadership in the dark. It was chaotic, exhausting, and ultimately unsustainable. Now, PagerDuty connects UWM's ServiceNow and Microsoft Teams into an automated loop. Incidents route to the right subject-matter expert through escalation policies. Engineers can then acknowledge and act directly from Teams without context-switching. "With PagerDuty, we get the right context to the right people who can fix the problem at the right time," says Jim Wallace, Operations Administrator at UWM. For Wallace, the change runs even deeper than speed: "[PagerDuty] is revolutionizing communication across IT." Because disruptions to UWM's broker-facing systems now get caught and resolved before they stall applications, loans keep moving. UWM now has an average close rate of 14 days - more than three weeks faster than the 38-day industry benchmark. Build toward autonomous resolution. As AI agents move into production, they introduce failure modes that legacy monitoring systems weren't built to catch. Hallucinations, model drift, and incorrect tool calls create subtle degradation in response quality. Without a system that treats AI quality drift as a high-priority incident, problems often reach customers (and impact business outcomes) before they reach the response team. Arize was built to close that gap. Their AI and agent engineering platform helps teams observe, evaluate, and improve AI agents in production by tracing every interaction, scoring outputs for quality, and tracking behavior over time. But detection is only part of the work. By integrating with PagerDuty, Arize can turn quality alerts into actionable guidance. Arize catches quality issues early to lower mean time to detection. PagerDuty gets each incident to the right person fast to lower mean time to resolution. When an Arize evaluation crosses a threshold, it fires an alert into PagerDuty, which routes it to the right responder with the context they need to triage. "Arize and PagerDuty together turn AI quality into a proactive operational discipline," says Richard Young, Technical Director of Partner Solutions Architecture at Arize. The bigger payoff is what Young calls "self-improving" agents. Rather than shipping once and slowly degrading, agents connected to Arize and PagerDuty continuously improve with use. Every interaction becomes data, and that data informs the next version. Turn raw signals into targeted action for the agentic era. The Warriors, UWM, and Arize had a problem many enterprises are facing right now: too much noise to act quickly. Across industries, PagerDuty is the layer that turns overwhelming raw signals into prioritized incidents and routes them to the right responder, with the right context. Problems get resolved upstream, before they reach the people who matter most. With PagerDuty, companies are turning fragmented data into intelligent action - and with every incident resolved, the platform gets smarter, compounding operational resilience over time. See how PagerDuty can help your team develop proactive incident response workflows. Schedule a demo today.
How Smartsheet's Data, AI, & Platform Engineering teams use Monte Carlo to catch issues before they reach the business. Overview. Smartsheet, the enterprise platform for modern work management, serves over 85% of Fortune 500 enterprises. It also runs a data and AI platform that powers internal analytics, ML models, and product intelligence at scale. As the volume and complexity of Smartsheet's data estate grew, so did the cost of data downtime: silent pipeline failures, freshness gaps, and schema drift that reached downstream dashboards before anyone noticed. Monte Carlo's agent trust platform became a central pillar of Smartsheet's reliability strategy by helping the team implement proactive, automated observability across data and AI. The challenge. Before Monte Carlo, Smartsheet's data engineering teams faced a set of compounding reliability problems common to any organization scaling its data estate aggressively: * Data issues reaching dashboards and business decisions before engineers could catch them * Manual, custom unit tests that were expensive to write and slow to catch schema changes or volume drops * Incident ownership spread across Slack threads with no structured workflow for assignment or resolution tracking * A growing blind spot in agentic and ML workflows, where LLM interactions were a "black box" for validation teams * Alert fatigue from misconfigured or overly sensitive monitors during initial setup phases As Dharmendra D., Senior Software Engineer on Smartsheet's Data & AI Platform team, described it: teams were handling incident coordination "in a much messier way across Slack threads" - with no clear ownership or resolution tracking to prevent the same issues from resurfacing. The solution: Monte Carlo's agent trust platform. Smartsheet deployed Monte Carlo as an end-to-end observability layer across their data and AI stack, integrating tightly with Databricks, Snowflake, Looker, PagerDuty, and Slack. The platform's automated ML-driven monitoring, field-level lineage, and incident management workflows replaced the manual, fragmented approach that had previously left teams reactive rather than proactive. Proactive anomaly detection before issues reach the business. Monte Carlo's ML models learn baseline behavior for each data asset and flag deviations - freshness drops, volume anomalies, schema changes - before downstream consumers notice. Rather than flooding teams with noise, the system surfaces only meaningful anomalies, "significantly reducing alert fatigue and helping our team focus on real issues rather than chasing false positives." Seamless integration with the modern data and AI stack. Smartsheet's engineers operate across a multi-tool environment spanning Databricks, Snowflake, Looker, PagerDuty, and Slack. Monte Carlo's broad integration surface made centralized observability possible without disrupting existing workflows. End-to-end lineage that cuts debugging time. Lineage visualization across data and AI assets has been a recurring need at Smartsheets, with engineers describing it as the feature that most directly translated to hours saved. The ability to trace an issue from a Looker dashboard all the way back to a Snowflake warehouse, or from a pipeline failure to its upstream source, eliminated the manual investigation that previously consumed debugging cycles. Being able to trace data from source to consumption in a clean, interactive graph saved Smartsheet engineers hours of investigation during incidents. Structured incident management replacing ad-hoc Slack coordination. One of the most operationally significant improvements across Smartsheet's teams was the shift from informal incident handling to structured, ownership-driven workflows. Monte Carlo's incident management module brought clear assignment, severity classification, and resolution tracking to what had previously been a coordination problem. "The incident management workflow is a highlight as well," noted Dharmendra D. "It keeps the team aligned on data quality issues with clear ownership and resolution tracking - something we previously handled in a much messier way across Slack threads." The result: fewer escalations and faster resolution across Smartsheet's Data & AI Platform. On ROI: "For a platform team, the ROI shows up as fewer escalations and faster incident resolution." The time saved debugging incidents, the reduction in manual monitoring effort, and improved organizational trust in data all compounded into measurable returns. Agent observability for emerging AI workloads. Smartsheet engineers are seeing the value of Monte Carlo beyond traditional data pipelines and into the agentic layer - a particularly resonant point given Smartsheet's active AI and ML development. Ruchir K. described the team's challenge: "In terms of Agent Observability, LLM interactions can be a bit of a black box for validation teams. We implemented an internal judge system for LLM-based projects, but Monte Carlo has also helped us get the big picture on how well our models are performing." Monte Carlo provided the visibility layer where internal tooling fell short. For teams managing multiple data squads under a larger analytics function - like Smartsheet's - Monte Carlo also simplified coverage tracking: "One of our ongoing challenges has been making sure all the different teams have proper coverage for our IP. We have a lot of squads under Analytics, and this has helped us keep the process moving so we can consistently ensure our products are covered appropriately." Results. Across Smartsheet deployments, Monte Carlo delivered on three fronts: engineering efficiency (eliminating manual unit test writing, cutting debugging time through field-level lineage, and freeing teams for higher-order work), data and AI reliability (catching issues before they reached dashboards, reducing alert fatigue through ML-calibrated detection, and building cross-organizational trust in data and agents), and operational maturity (replacing ad-hoc Slack coordination with structured incident ownership, clear resolution tracking, and scalable multi-team coverage). About Smartsheet. Smartsheet (NYSE: SMAR) is the enterprise platform for modern work management, helping organizations plan, capture, manage, automate, and report on work. Over 85% of Fortune 500 companies trust Smartsheet. Headquartered in Bellevue, WA. Its promise: MonteCarlo will show you the product.
PagerDuty announces Arnaud Lagarde, vice president of EMEA. SAN FRANCISCO-(BUSINESS WIRE)- PagerDuty, Inc. (NYSE: PD), a leader in AI-first operations management, today announced the appointment of Arnaud Lagarde as vice president of EMEA. Lagarde will lead PagerDuty's next phase of growth in the EMEA region, bringing the entire incident management lifecycle to customers across EMEA to solve their biggest digital challenges. "We are thrilled to appoint Arnaud as vice president of EMEA, since he brings a wealth of enterprise sales relationships and years of experience growing this region," said Todd McNabb, chief revenue officer at PagerDuty. "Arnaud brings a specific combination of deep technical expertise and leadership that will be critical for PagerDuty's customers, partners and employees. He is a great fit for PagerDuty and we look forward to his impact." Lagarde brings to the role over 20 years of experience spanning companies like Automation Anywhere, CA Technologies and BMC. Over the past two decades, he has worked closely with founders, investors and executive teams to establish product-market fit, develop enterprise go-to-market strategies and build high-performing revenue organizations. He brings extensive sales and domain expertise in AI-driven companies and has a strong track record for building high-performing regional teams and executing go-to-market strategies. "PagerDuty has a massive opportunity across EMEA at a critical turning point for AI-first digital operations," said Arnaud Lagarde, vice president of EMEA, PagerDuty. "With my experience industrializing AI at scale in mission-critical processes and regulated industries, I look forward to working with our regional customers to unlock entirely new levels of operational efficiency. Moving forward, my priority is to empower our teams and ecosystem partners to accelerate growth, drive measurable business outcomes, and foster enduring enterprise relationships." Lagarde received his Master of Art in Business Studies from ESIC ESSICA Bordeaux and graduated from IUT de Bordeaux earning a degree in business studies. About PagerDuty Inc. PagerDuty, Inc. (NYSE: PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing, and remediating issues, the PagerDuty Operations Cloud acts as the central control plane for the modern enterprise - orchestrating AI agents and automated workflows with context from over 750 integrations. Trusted by approximately two-thirds of the Fortune 100 and nearly half of the Fortune 500, PagerDuty is the industry standard for organizations scaling resilient, autonomous operations. Learn more and try it for free at www.pagerduty.com. The PagerDuty Operations Cloud The PagerDuty Operations Cloud is an AI-powered platform that automates and orchestrates the entire incident management lifecycle - from detection to resolution, providing resilience at scale. Designed for mission-critical operations, the platform empowers teams to identify and diagnose disruptions in real time, mobilizing the right teams to quickly streamline workflows to solve digital issues before they become incidents. The PagerDuty Operations Cloud is essential for delivering flawless, always-on digital experiences that organizations and consumers expect today.
AI Orchestrations: Your easy button for proactive operations. This blog post is part of PagerDuty's ongoing series on how PagerDuty, Inc. is helping customers navigate their journey towards autonomous operations. Read on to learn about how AI Orchestrations builds towards this vision. "We should automate this." Sound familiar? For many operations teams, that sentence never becomes action. Building event orchestration rules demands deep platform expertise, time no one has, and the ability to spot which patterns in your data actually matter. So alert fatigue lingers. Toil piles up. Automation stays on the distant roadmap. PagerDuty's new AI Orchestrations capability was built to close that gap - for good. From reactive to proactive, automatically. AI Orchestrations analyzes your historical event and incident data (including how responders have manually intervened during past incidents) and surfaces ready-to-apply recommendations for event orchestration rules. No rule syntax. No expert configuration. Just a plain-language description of what the rule does, how many alerts or incidents it would have improved, and a confidence score based on past accuracy. The result is a global recommendations view across all of your services: every AI-generated suggestion, ranked, filterable by team, and actionable in a single click. What gets recommended. The system surfaces several types of rules based on detected patterns in your data: * Suppression - silence low-value alerts before they become incidents * Severity - dynamically set alert severity so responders know what demands immediate attention * Priority - automatically triage incidents so teams focus on what matters Each recommendation includes impact metrics - how many past events would have been affected - and a precision score so teams can validate before applying. Human in the loop, friction removed. AI-suggested rules live in a dedicated AI Orchestration Segment that runs after your existing manual rules. Your current logic always runs first. When a recommendation looks right, you apply it with one click. When it's not relevant - or when you'd rather fix the underlying monitor than suppress the alert - you dismiss it, and that feedback improves future recommendations for your service. This keeps humans in control while eliminating the #1 barrier to automation adoption: the expertise required to build rules in the first place. The flywheel effect. Every incident your team responds to contains signal. AI Orchestrations turns that signal into automation - and better automation means fewer incidents, less noise, and faster resolution. Teams using event-driven automation see 46% fewer incidents and up to 30% lower FTE costs. Over time, the platform gets smarter as your data evolves, creating a continuous flywheel: incidents inform automation, automation prevents incidents. This is how PagerDuty delivers on the vision of Autonomous Operations - not by replacing human judgment, but by making it easier to apply at scale. Available now for AIOps customers. AI Orchestrations is Generally Available today for PagerDuty AIOps customers. If your team is spending time manually triaging repetitive alerts, struggling with low event orchestration adoption, or onboarding new members who don't have bandwidth to configure automations - this is the place to start.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
IPO
Headquarters
San Francisco, California
Founded
2009
Find jobs on Simplify and start your career today