
Work Here?
Arize AI provides a platform for AI observability and LLM evaluation that helps teams monitor, troubleshoot, and improve the performance of machine learning models, including generative models, NLP, computer vision, and recommender systems. It works by delivering analytics and workflows, tracing LLM operations (traces and spans), retrieval augmented generation (RAG), and task-based evaluations for issues like hallucinations, relevance, and citation checks. The platform distinguishes itself by offering end-to-end AI observability and LLM evaluation for top AI companies, with embedding-based RAG analysis to diagnose missing or irrelevant context. The goal is to help AI teams ensure models perform optimally and continuously improve their AI systems.
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
201-500
Company Stage
Series C
Total Funding
$131M
Headquarters
Berkeley, California
Founded
2020
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$131M
Above
Industry Average
Funded Over
4 Rounds
Industry standards
Health Insurance
Dental Insurance
Vision Insurance
401(k) Retirement Plan
Unlimited Paid Time Off
Parental Leave
Mental Health Support
Flexible Work Hours
From Signal to PR: What if your agents got better every time they failed? Published july 29, 2026. Co-Authored by Chris Cooning, Head of Product Marketing & Sally-Ann DeLucia, Head of Product & Jason Lopatecki, Co-founder and CEO & Aparna Dhinakaran, Co-founder & Chief Product Officer. What if your agents got better every time they failed? It's 2 AM, and the pager goes off. User frustration rates are climbing, so someone opens a laptop, and the clock that matters (time to fix) has barely started. The eventual patch may be two lines, but the expensive part comes first: finding the right traces, reconstructing the failure, forming a theory, and locating the responsible code. Arize is launching Signal, a managed agent built into Arize AX that takes on that investigative work continuously. Signal reviews production traces, identifies recurring failure patterns, and groups them into ranked issues. Each issue includes supporting evidence, a root-cause analysis, and a proposed fix. Signal doesn't stop at surfacing issues. With a repository connected, a managed agent can carry the investigation into the codebase, propose a fix, and open a pull request, turning production telemetry into a reviewable change. This launch points toward a more ambitious future. Arize believe agents will not only run in production, but also help improve the systems they run in. In fact, Signal is part of the agent improvement loop in Arize AX where production behavior becomes evidence, evidence becomes an investigation, and the investigation becomes a reviewable change. The engineer still makes the call; they just begin with a diagnosis instead of a blank query box. Tl;dr. * Signal finds the issue. It continuously reviews production traces, groups related failures, and surfaces the evidence, likely cause, and proposed fix. * Managed Agents help close the loop. With repository access, they can inspect the relevant code, propose a patch, and open a pull request. * Agent Studio makes the loop configurable. Teams can run guided or custom investigations once, on a schedule, or in response to an operational trigger. * Humans remain in control. The agent investigates and proposes; the engineer reviews and ships. Signal ships on Arize AX Free and Pro today. Full managed agents, including Agent Studio, presets, and repository access, are available to Enterprise customers in beta. Observability has a new reader. For decades, production telemetry had one real consumer: a human. Applications emitted logs, metrics, and traces, and dashboards organized them. Alerts woke someone up, and an engineer then translated that telemetry into an explanation. Coding agents accelerated the final step. Once a developer understood the problem, an agent could write the patch for them. But the human still had to consume the telemetry, isolate the failure, and turn the investigation into a prompt. Signal moves the agent upstream. This is the next step toward self-improving software: the telemetry you capture stops being something only a human reads and becomes an input an agent can immediately act on. From traces to action. * Evidence comes from traces and evaluations. Traces capture what the agent did: model calls, retrievals, tool executions, inputs, outputs, timing, and errors. Evals help distinguish a technically successful run from a good result. * Context connects runtime behavior to its cause. That can include logs, application spans, metadata, and, when connected, the repository where the change needs to land. * A trigger determines when the investigation runs. Signal continuously sweeps new production traces, while broader managed-agent workflows can run on a schedule or in response to an operational event. Together, those pieces create an improvement loop: * Investigate * Propose * Review * Ship * Observe again This makes observability, especially tracing, more important. Traces serve as the source of truth and the backbone of the feedback loop. The higher the quality of the traces, the better the agent can investigate failures and identify the right code to change. Paired with strong evaluations that distinguish error modes from silent quality issues, they help the agent turn "this behavior is broken" into "this is the code that should change." Start with Signal, extend with Managed Agents. Signal comes preconfigured for continuous production reliability. Turn it on for a tracing project, and it begins reviewing new runs, tracking issues it has already seen, and surfacing emerging failure patterns. For teams that want to go beyond that workflow, Managed Agents and Agent Studio in Arize AX include templates for investigating failing traces, triaging monitor alerts, analyzing costs, running recurring health checks, inspecting repositories, and proposing pull requests. Teams can also start with a blank agent and build around virtually any engineering workflow they need. Signal is not the endpoint, though. It is an early piece of a broader shift toward self-improving agent systems: software that can observe its own behavior, identify where it is failing, and propose the next improvement for a human to approve. The future is not an agent rewriting itself in production. It is a controlled loop in which every trace can become evidence, every failure can become an investigation, and every investigation can become a reviewable change. Signal is available in Arize AX today. Turn it on against the traces and evaluations you already capture.
How PagerDuty powers intelligent incident response for the Golden State Warriors, UWM, and Arize. The best incident is the one your customers never notice. A payment clears, a stream doesn't buffer, a loan closes on time - all because upstream, your team diagnosed and resolved the problem before it spread. But that's getting harder to pull off. PagerDuty is already ahead of that curve - and customers are seeing compounding results. With 59% of organizations already actively incorporating AI into operational workflows, applications are generating more signals (and more novel failure modes) than legacy monitoring systems were ever designed to handle. That means more noise for your team to sift through to reach the root of the issue. Meanwhile, customers get frustrated, and you creep closer to the limits of your SLAs. The enterprises winning in this era aren't just collecting more signals. They're making them actionable. PagerDuty's Operations Cloud uses AI Ops - including intelligent triage, automated runbooks, and the PagerDuty SRE Agent - to accelerate alert triage, route incidents to the right responders with full context, and resolve routine issues automatically. Now, teams can detect, fix, and even prevent incidents before they impact end users. Here's what the benefits look like for three PagerDuty customers. Catch critical issues before they affect service. The Golden State Warriors run one of the most technology-forward operations in sports. At Chase Center, their home arena in the Bay Area, the fan experience spans online ticket sales, in-venue food and beverage transactions, and a website that serves more than 70,000 active users a month. Fans expect a seamless experience, so every system has to perform perfectly during the game, when there's no option to delay service. Before PagerDuty, executives or fans flagged 80% of issues before the Warriors' IT team ever knew about them. In one case, a bug buried in third-party code slipped through testing and reversed an entire category of transactions before anyone caught it the next day. With PagerDuty, their system now tracks logs across the Warriors' digital platforms, surfaces irregular patterns, and alerts the team the moment something looks off - sending them to the people who can fix them before they cascade into a game day service failure. PagerDuty also helped the team map its incident response from end to end, defining escalation paths for both critical and non-urgent incidents. Each alert now gets prioritized, so the team applies the right level of response every time. "We need to be able to respond quickly, and PagerDuty allows us to inform the right person at the right time so that we're able to intervene and get a resolution as fast as we can," says Nick Manning, Senior Director of Consumer Product and Emerging Technologies at the Golden State Warriors. Now the team catches issues before executives or fans do, maintaining a seamless fan experience instead of scrambling to react. Modernize systems for better communication. United Wholesale Mortgage (UWM) is ranked among the nation's top mortgage lenders. When their service is disrupted, borrowers aren't just frustrated - they're delayed from closing on their homes. But for years, UWM's incident response workflows actually got in the way of their mission. Before PagerDuty, brokers reported issues to the help desk, who then tracked down a few key people by email or text, then tagged the rest of the IT floor in Microsoft Teams (day or night). Each manual handoff burned critical minutes and fragmented context, leaving leadership in the dark. It was chaotic, exhausting, and ultimately unsustainable. Now, PagerDuty connects UWM's ServiceNow and Microsoft Teams into an automated loop. Incidents route to the right subject-matter expert through escalation policies. Engineers can then acknowledge and act directly from Teams without context-switching. "With PagerDuty, we get the right context to the right people who can fix the problem at the right time," says Jim Wallace, Operations Administrator at UWM. For Wallace, the change runs even deeper than speed: "[PagerDuty] is revolutionizing communication across IT." Because disruptions to UWM's broker-facing systems now get caught and resolved before they stall applications, loans keep moving. UWM now has an average close rate of 14 days - more than three weeks faster than the 38-day industry benchmark. Build toward autonomous resolution. As AI agents move into production, they introduce failure modes that legacy monitoring systems weren't built to catch. Hallucinations, model drift, and incorrect tool calls create subtle degradation in response quality. Without a system that treats AI quality drift as a high-priority incident, problems often reach customers (and impact business outcomes) before they reach the response team. Arize was built to close that gap. Their AI and agent engineering platform helps teams observe, evaluate, and improve AI agents in production by tracing every interaction, scoring outputs for quality, and tracking behavior over time. But detection is only part of the work. By integrating with PagerDuty, Arize can turn quality alerts into actionable guidance. Arize catches quality issues early to lower mean time to detection. PagerDuty gets each incident to the right person fast to lower mean time to resolution. When an Arize evaluation crosses a threshold, it fires an alert into PagerDuty, which routes it to the right responder with the context they need to triage. "Arize and PagerDuty together turn AI quality into a proactive operational discipline," says Richard Young, Technical Director of Partner Solutions Architecture at Arize. The bigger payoff is what Young calls "self-improving" agents. Rather than shipping once and slowly degrading, agents connected to Arize and PagerDuty continuously improve with use. Every interaction becomes data, and that data informs the next version. Turn raw signals into targeted action for the agentic era. The Warriors, UWM, and Arize had a problem many enterprises are facing right now: too much noise to act quickly. Across industries, PagerDuty is the layer that turns overwhelming raw signals into prioritized incidents and routes them to the right responder, with the right context. Problems get resolved upstream, before they reach the people who matter most. With PagerDuty, companies are turning fragmented data into intelligent action - and with every incident resolved, the platform gets smarter, compounding operational resilience over time. See how PagerDuty can help your team develop proactive incident response workflows. Schedule a demo today.
Arize AI: production-ready AI agent observability workshops. 4h ago · 0:00 listen · Source: TipRanks Summary. Arize AI is highlighting two hands-on workshops at the AI Engineer World's Fair. These sessions focus on moving from experimental to production-ready AI agents. Laurie Voss will lead workshops that cover tracing, evaluations, experiments, and production monitoring. A financial-analyst agent will be used as a practical example. The workshops include a 101-level session on instrumentation, error analysis, and feedback loops. A 201-level workshop extends to session-level evaluations, RAG quality scoring, and autonomous issue investigation using Arize's Signal product. This emphasis on end-to-end observability suggests Arize AI is positioning itself as a key infrastructure provider for production AI systems. This could be particularly relevant for data-intensive sectors like financial services. These workshops aim to deepen developer engagement and could enhance Arize AI's standing in the AI observability market. This is an AI-generated audio summary. Always check the original source for complete reporting.
How Arize AI scaled Phoenix to millions of downloads. Sat Apr 25 2026 * Challenge: Large language models hallucinated unpredictably in production and developers lacked standard tools to evaluate and monitor them. * Solution: Arize AI launched Phoenix, an open source observability and evaluation library specifically for AI applications. * Results: Phoenix surged to over 2 million monthly downloads and Arize AI secured a $70 million Series C round in 2025. * Investment/Strategy: Betting completely on an open source, product led distribution model to become the default standard before monetizing at the enterprise level. The problem. Deploying AI models in production used to be a massive blind spot. Engineering teams would spend millions training models or fine tuning prompts, only to watch them fail spectacularly when exposed to real users. Hallucinations, data drift, and unexpected toxicity were rampant. The standard software monitoring tools built for traditional applications were completely useless when applied to the probabilistic nature of large language models. Founders and developers were forced to string together custom logging scripts and manual evaluation spreadsheets. They had no systematic way to trace complex multi step agent workflows or compare how different models handled edge cases. This "last mile" of deployment became a massive bottleneck. The entire industry was moving fast, but teams hesitated to push generative AI features to production because they simply could not measure if the output was safe or accurate at scale. The execution & GTM strategy. The product led distribution strategy. Arize AI realized that trying to sell complex enterprise software directly to executives for an emerging technology was the wrong move. Instead, they built Phoenix as an open source library and gave it directly to the engineers feeling the pain. By making the core tracing and evaluation tools free and frictionless, Phoenix became the default infrastructure for developers experimenting with LLMs. Developers could run Phoenix locally to visualize their prompt executions without navigating a procurement process. The technical moat. The core technical advantage of Phoenix was its adoption of OpenTelemetry standards for tracing. Rather than forcing engineers to learn a proprietary logging format, Arize AI built Phoenix to plug into existing open standards. This mechanism allowed seamless integration with frameworks like LangChain and LlamaIndex. When developers instrumented their applications with these popular frameworks, Phoenix was often the most logical and native way to visualize the trace data. The interoperability itself became a massive technical moat. The monetization layer. While Phoenix served as the massive top of funnel acquisition engine, Arize AI structured their business model around the complexities of scale. The open source tool was perfect for local development and small projects. However, when large enterprises needed to manage role based access control, host massive evaluation datasets, and run continuous monitoring pipelines across thousands of concurrent users, they needed a managed service. Arize AI effectively monetized the heavy infrastructure demands of enterprise scale while keeping the individual developer experience completely free. The Results & takeaways. * Surpassed 2 million monthly downloads for the Phoenix open source library. * Raised a $70 million Series C round in early 2025, bringing total funding to over $131 million. * Expanded strategic partnerships with major cloud providers like Microsoft Azure. * Established the industry standard for LLM evaluation and observability. What a small startup can take from them: If you are building infrastructure for a completely new developer paradigm, do not hide your core value behind a paywall. By open sourcing their core tracing engine, Arize AI embedded themselves into the developer workflow before their competitors even got a sales meeting. Build the tool that developers use locally, and eventually, their companies will pay you to host it globally. Frequently asked questions. Arize AI relied heavily on open source product led growth. They distributed Phoenix for free to individual developers to build a massive user base, which eventually converted into enterprise contracts as those developers scaled their applications.
Arize AI, Inc is excited to introduce Alyx, the next evolution in Arize's intelligent assistant.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
AI & Machine Learning
Company Size
201-500
Company Stage
Series C
Total Funding
$131M
Headquarters
Berkeley, California
Founded
2020
Find jobs on Simplify and start your career today