Full-Time
AI SRE platform for autonomous remediation
$150k - $300k/yr
New York, NY, USA
In Person
Five days in-office per week in New York.
See people who can refer or advise you
Traversal provides an AI-powered platform for site reliability engineering and observability. Its AI SRE agent autonomously detects, troubleshoots, and resolves production incidents by analyzing telemetry and performing root-cause analysis to identify underlying causes. It combines large language models with causal machine learning to orchestrate real-time remediation and offers proactive health checks. It can be deployed as a standalone product or as an intelligence layer on existing observability stacks, including on-premise hosting, to serve enterprises like cloud providers and large SaaS firms, with a goal of reducing downtime and moving systems toward self-healing.
Company Size
51-200
Company Stage
Series A
Total Funding
$48M
Headquarters
New York City, New York
Founded
2023
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Health Insurance
Flexible Work Hours
Company Equity
PepsiCo taps Traversal to build an agentic mission control. TABLE OF CONTENTS At a glance. Global food and beverage leader PepsiCo operates one of the world's most complex supply chains, with over 3000 applications supporting manufacturing, distribution, and customer operations across multiple geographies. Its OnePepsiCo Operations Centre (OnePOC) team - the first line of defense for production incidents - faced a critical operational challenge: over 80% of Pepsi's major incidents were first detected by employees on the warehouse floor or frontline staff, not by their internal teams. Traversal partnered with PepsiCo to transform their operational command center from a reactive escalation point into a proactive incident prevention engine: a step towards agentic mission control. By automatically triaging and ranking thousands of alerts based on business-context analysis, mapping system interdependencies, and delivering root cause analysis within minutes when incidents occur, Traversal empowered PepsiCo to preemptively identify major incidents and coordinate faster, highly targeted incident resolution across vendor and engineering teams. This resulted in a more than 200% increase in alert processing capacity, freeing up over 500 engineering hours per month across the subset of applications where Traversal was deployed. The challenge. PepsiCo's One PepsiCo Operations Centre (OnePOC) is responsible for safeguarding business continuity across a global, highly interconnected supply chain. Its mandate is not simply to manage alerts, but to detect issues early, coordinate cross-functional response, and minimize operational disruption before it reaches frontline teams. However, operating at PepsiCo's scale introduced structural challenges that compounded across the incident lifecycle: * Enterprise-Scale Signal Complexity: PepsiCo's digital ecosystem generates over 500,000 alerts each month. While only a small fraction correspond to business-impacting incidents, each signal must be evaluated to determine relevance and severity. This volume resulted in sustained alert backlogs - including tens of thousands of unresolved alerts and hundreds flagged as high priority - creating systemic friction in triage and prioritization workflows. The challenge was not organizational efficiency, but the sheer scale and signal-to-noise ratio inherent in operating one of the world's largest supply chains. * Escalation Across a Distributed Operating Model: When alerts matured into Major Incidents (MIMs), resolution required coordination across multiple internal teams, infrastructure partners, and application owners. Given the distributed and interdependent nature of PepsiCo's environment, isolating root causes often involved cross-domain expertise and time-zone coordination. This extended investigation windows and increased business exposure during high-impact events. * Fragmented Infrastructure Intelligence: PepsiCo's observability landscape spans multiple specialized platforms, each providing valuable telemetry but requiring manual correlation to build a complete operational picture. Even understanding the blast radius of an issue required navigating disparate tools and query languages. The limitation was not a lack of data - it was the absence of unified, contextualized infrastructure intelligence aligned to business impact. Its deployment. PepsiCo's pilot phase focused on proving value across 4 core business-critical applications which tied most directly to revenue generating operations. Traversal integrated with PepsiCo's observability landscape, connecting via read-only access to six unique data sources - a mix of legacy and modern platforms: Elastic, AppDynamics, ServiceNow, Azure Data Lake (ADLS), ThousandEyes, and Grafana - all via a SaaS deployment model. The 4-week pilot period was designed with rigorous metrics across three key success criteria: RCA accuracy on both historical and live incidents & alerts, quantifiable MTTR reduction through back-testing, and direct user feedback validation. Over the pilot period, these capabilities demonstrated measurable impact - over 80% RCA accuracy, 500+ monthly engineering hours saved, and elimination of a backlog of 700+ high-severity alerts which helped prevent incidents. Strong user engagement across both OnePOC and MIM teams, validated through direct feedback and careful product analytics, led to approval for strategic expansion from the initial 4 pilot applications to a planned rollout across 66 applications. Traversal's Impact at PepsiCo. Traversal fundamentally transforming how PepsiCo addresses each of its core operational challenges: * Autonomous Alert Triage and Prioritization: In September 2025, Traversal automatically handled thousands of PepsiCo's 500,000 monthly alert volume. With ~10 minutes saved per high-priority alert and at over 100 high-priority alerts each day, Traversal unlocked over 500 engineering hours saved per month, while eliminating the backlog of 700+ open high-severity alerts that required immediate attention. When alerts are now escalated to an engineer, Traversal provides full context - specific infrastructure nodes, root cause hypotheses, and business impact - eliminating reliance on tribal knowledge and enabling faster, informed responses. * Agentic RCA for MIMs: With over 80% RCA accuracy across incidents, Traversal's incident RCA product significantly reduced the time-consuming investigation phase that made incident response costly. Root causes that previously required hours of coordinating across teams are now identified in minutes, reducing downtime and freeing engineering capacity. At PepsiCo, this has translated to an average of 500 engineering hours saved monthly, even with limited scope. * Single Pane of Glass for Infrastructure Intelligence: Traversal provides a single pane of glass across Pepsi's full infrastructure stack, traversing services, dependencies, and telemetry across all observability systems - rapidly retrieving and correlating data to answer questions like "Which components are affected by this incident?" in minutes. This gives OnePOC members and engineers the operational context and response speed typically associated with senior SREs, shifting how PepsiCo's operational teams access critical infrastructure knowledge. Instead of engineers struggling to navigate PepsiCo's complex observability landscape, Traversal enabled natural language queries to interface in a unified manner across all their data sources. Underlying these capabilities is Traversal's Production World Model(TM), which compresses and re-indexes PepsiCo's fragmented observability, dependency, code, tribal knowledge, and infrastructure data into a unified, machine-readable model of the entire production environment. Traversal's Causal Search Engine(TM) then investigates over it, testing thousands of hypotheses in parallel to arrive at a causally consistent diagnosis of likely root cause, affected services, and business impact. Traversal has helped across all of PepsiCo's mission control operations and is deeply embedded in daily workflows. The business impact is substantial: investigation time has been reduced by 70%, enabling faster resolution and materially reducing operational friction during incidents, freeing up engineering capacity. Inside a real incident. At 4:06AM on October 31, 2025, warehouse workers in one of PepsiCo's main distribution centers began experiencing failures with a critical application: the technology that tells workers which pallets to move and where. When warehouses can't move inventory, distribution capacity drops, and manufacturing plants lose storage space for finished goods. The AppDynamics alert triggered immediately. Traversal automatically launched an investigation. Within 3 minutes, it identified the root cause: CPU saturation on a shared middleware node, correlated with a critical application deployment 21 minutes earlier. Traversal mapped the impact to co-hosted services, assessed business criticality, and recommended response escalation to the Platform Operations team. OnePOC escalated with full context, and PepsiCo's infrastructure partner knew exactly which host to inspect and what to fix. The incident was contained within 30 minutes-something that would've taken hours for the initial manual investigation alone. Towards an agentic mission control. PepsiCo's deployment of Traversal represents more than operational efficiency gains: it's a fundamental architectural and conceptual shift towards an agentic mission control. Building this future requires solving all three of PepsiCo's core operational challenges - and Traversal addresses each one. Alert triage, incident RCA, and seamlessly understanding your infrastructure, which once required extensive manual effort now occur autonomously with minimal human intervention. Hence, enabling a small team of human orchestrators to efficiently coordinate complex multi-team responses, and empowering senior engineers to focus on system design and long-term resilience rather than reactive firefighting. As the system learns from resolution patterns, the goal is to move towards self-healing where common issues are automatically remediated before human intervention is needed. To see Traversal in action, book a demo today. "Operating at PepsiCo's scale - 3,000+ applications, hundreds of thousands of alerts - requires intelligent automation beyond traditional monitoring. Traversal's AI SRE agents cut through this enormous complexity, automatically triaging alerts and surfacing root causes in minutes rather than hours. What previously required coordinating 20+ engineers across time zones now happens autonomously. It's transforming our operations from reactive firefighting to preventing incidents before they reach the warehouse floor - the future of reliability at global scale." Vinod Chilakalapudi - Director of IT Operations, PepsiCo
Traversal, an AI lab building agents for enterprise site reliability engineering, has announced six senior leadership hires across go-to-market and engineering within a single month. The company's headcount has grown to over 90, representing a 110 per cent increase in six months. New appointments include Jim Cavanaugh as SVP of Worldwide Sales, Ryan Powers as SVP of Marketing, Patrick Wade as VP of Worldwide Field Engineering, and Maxime Petazzoni as Head of Engineering. The hires bring experience from companies including Cribl, Redis, SignalFx and Splunk. The expansion follows Traversal's recent investment from Amex Ventures and deployment across American Express. A Fortune 100 financial services case study showed 32 per cent reduction in potential mean time to resolution and 82 per cent root cause analysis accuracy.
Traversal, the frontier lab building AI agents for enterprise-grade site reliability engineering (SRE), today announced a strategic investment from Amex Vent...
American Express has partnered with and invested $5 million through Amex Ventures in Traversal, an AI-driven site reliability engineering startup founded by researchers from MIT, Columbia and Cornell. The credit card company will deploy Traversal's platform across its global technology infrastructure. Traversal uses large language models, AI agents and causal machine learning to analyse operational telemetry data across multiple monitoring systems, helping diagnose and resolve technology outages more quickly. The platform aims to automate work traditionally requiring dozens of engineers collaborating in "war rooms" during incidents. The startup has raised approximately $53 million to date. Its technology addresses fragmentation in the observability market by inferring cause-and-effect relationships across different monitoring platforms, moving beyond simple pattern detection to root cause analysis.
Cloudways launches self-healing site reliability solution, powered by Traversal. At a glance. Cloudways, a leading managed cloud hosting platform, partnered with Traversal to transform its customer support and site reliability experience. Powered by Traversal's AI SRE platform, Cloudways Copilot is an end-to-end self-healing solution that enables users to identify issues and remediate them instantly with a single click. This is the first instance of self-serve site reliability as a service. Following strong adoption and positive feedback, Cloudways Copilot entered into general availability in August 2025, rolling out its issue diagnostics and self-healing solution to all 845k+ customer applications. The challenge. Cloudways - recently ranked by CNET as the number one web hosting software for developers - serves as the cloud infrastructure management platform for website hosting for digital agencies, developers, and small businesses across the globe. Like any platform that is mission critical for a diverse customer base with a broad range of technical needs, Cloudways requires a strong, responsive support workflow to ensure reliability at scale. Prior to partnering with Traversal, Cloudways customers facing issues like slow site performance, failing service, or DDoS attacks, would report their problem via chat or a helpdesk ticket, and receive diagnostic commands from a support engineer. Customers would attempt to run those commands themselves and, if unsuccessful, request remote assistance. The process often involved multiple back-and-forths and long delays in resolution due to customers' varying levels of technical expertise. To improve this experience, Cloudways partnered with Traversal to build an AI SRE with the ambitious goal of not just being a copilot for troubleshooting incidents, but an end-to-end autonomous troubleshooting and self-healing tool to over 845k applications hosted on the platform. Its deployment. Traversal began as a pilot with 500 Cloudways WordPress customers. For data privacy, troubleshooting for Cloudways customers required Traversal to access machine-level logs and metrics directly, rather than reading from a centralized observability stack. Traversal AI connected with custom Cloudways endpoints - for example, Sensu for alerts and Ansible for workflows - all via a custom proxy to meet enterprise-grade guardrails, reliability, and security standards. The resulting solution was launched in private preview as Cloudways Copilot, powered by DigitalOcean's proprietary Gradient AI platform. Its capabilities would include ingesting customer context, identifying the root cause of issues, and return recommended next steps for remediation - often within minutes. As confidence in Copilot's root cause identification grew, customers began asking for a way to apply fixes automatically. In response, Traversal Inc. launched a "SmartFix" feature, enabling users to automatically execute recommended remediations directly from the support flow with the click of a button. Cloudways Copilot is now in General Availability and is being rolled out to all Cloudways customer applications. It is currently performing over 1,000 investigations per day, with volume expected to grow to as many as 4,000 investigations per day as rollout completes. Traversal's impact at Cloudways. Cloudways Copilot constantly monitors the web stack, disk, inodes, and host health, detecting issues within seconds - from high-traffic anomalies like bot crawling and DDoS to system-level issues such as disk space exhaustion, inodes full, and service failures. It quickly analyzes the root cause and delivers clear, actionable recommendations, with the option to remediate automatically. This near-instant diagnosis helps recover optimal server performance with minimal effort, saving customers hours of manual troubleshooting. "We partnered with Traversal to build an end-to-end self-healing system - from alert to remediation. With over 95% accuracy, we can for the first time enable self-service reliability for our thousands of customers, instead of hours of frustrating back-and-forth with support - potentially saving millions in downtime and SRE costs." - Suhaib Zaheer, SVP & GM of Managed Hosting, Cloudways "With Copilot monitoring our servers and 47 applications, we identify problems before clients even experience issues - like getting automated insights that pinpoint exactly which applications are causing problems." "Cloudways Copilot & AI is a game-changer for reducing the amount of time spent taking care of your web server. It is the first good implementation of AI I've seen in a web host that actually makes my life as an agency owner easier." "Cloudways Copilot has transformed how we manage 180+ sites, saving our team 15 hours in just the last month. Instead of spending hours debugging, we now get detailed breakdowns that help us quickly resolve problems." Inside a real incident. At 2:07 PM, a WordPress site hosted by a web development company managing hundreds of sites on Cloudways began to slow down. Pages were timing out, CPU usage spiked, and some users saw 502 and 524 errors, but the root cause wasn't immediately clear. Normally, Cloudways Support would step in on behalf of the customer - spending 60 - 90 minutes collecting logs, isolating the issue, and coordinating with engineers. This time, the alert was handled by Traversal's AI SRE, streamlining the response without any manual triage: * 2:08 PM - Traversal began investigating on behalf of the customer. * 2:10 PM - It identified a set of abusive IPs overwhelming the site and outlined the root cause. * 2:12 PM - It proposed a self-healing action: block the malicious IPs and restart affected services, with UI-guided steps and a full remediation summary. * 2:13 PM - With a single click, the issue was resolved - end to end, in under 5 minutes. What would've taken hours was handled autonomously by Traversal, enabling Cloudways to respond to customer issues faster and more reliably - without manual triage or escalation.