
Work Here?
Work Here?
Work Here?
Grafana Labs builds observability and monitoring tools for cloud infrastructure and applications. Its flagship Grafana dashboard lets users visualize data from many sources in real time and set up alerts, with additional options like Grafana Enterprise for large deployments and Grafana Cloud as a managed service. The core open-source platform is complemented by commercial features and services that provide security, scalability, and dedicated support, appealing to both individual developers and large organizations. The goal is to help businesses keep digital services reliable and efficient by delivering scalable, real-time visibility into software and infrastructure.
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
Series D
Total Funding
$805.2M
Headquarters
New York City, New York
Founded
2014
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Total Funding
$805.2M
Above
Industry Average
Funded Over
6 Rounds
Industry standards
30 days of paid vacation each year on top of national holidays, parental leave, & sick leave
Health coverage
4% contribution match on our 401(k)
$1,500 learning and development stipend
Udemy subscription
Complimentary subscription to Headspace
Discounts on a wide variety of services, including entertainment, food, and fitness.
Remote Work Option
Global Employee Assistance Program
Identify anomalies, outlier detection, forecasting: How Grafana Cloud uses AI to make observability easier. 2026-08-27 - 6 min Note: This blog originally published in July 2024 and was updated in August 2026 to reflect the latest AI solutions and workflows in Grafana Cloud. Tech stacks keep getting more complicated, which means they're also more difficult to monitor. At Grafana Labs, Grafana is building AI solutions to help you understand these complex systems with less toil, so you can get to answers faster when something looks off. A big part of that is helping you catch anomalies earlier and pinpoint root causes, without needing to be an expert in every query language. In this post, Grafana'll explore a few of the AI-powered capabilities in Grafana Cloud designed to do just that, from forecasting and outlier detection to easier data exploration and incident response with Grafana Assistant, Watchers, and Investigations. Find anomalies: Agentic alerting, forecasting, and outlier detection. There are lots of ways AI can help you spot potential issues earlier in Grafana Cloud. Here are a few capabilities that help you identify when your system deviates from what's expected, whether that's against your own definition of "normal," against historical patterns, or compared with similar services and instances. Agentic alerting for your systems: Assistant Watchers. What if you know what you want to keep an eye on, but you don't know exactly which metrics or thresholds to alert on? That's why Grafana built Assistant Watchers. A capability of Grafana Assistant, the AI-powered agent in Grafana Cloud (more on that below), Watchers are always-on agents that check a scoped slice of your telemetry on a schedule. You describe what "normal" looks like in your own words - the service, the failure mode, the noise that's expected during a deployment. Assistant Watchers calibrates concrete checks against your metrics and logs, then raises concern when a run finds something notable. For example, let's say you deployed a new service and have some telemetry, but you aren't sure what to alert on yet. You can use a watcher to monitor the service and tune what "normal" looks like as you go. A watcher can stay quiet when signals look healthy, notify via Slack or a webhook when it flags a warning, and launch Assistant Investigations (again, more on that below) for critical findings. Use Watchers for recurring service health checks, noisy-but-important failure modes, and follow-up monitoring after an incident or deployment. Forecasting and outlier detection. When you already know which metrics matter, forecasting and outlier detection features in Grafana Cloud give you more targeted ways to identify unusual behavior. These tried-and-true approaches learn what to expect from your metrics over time or across groups of similar services, so you can detect deviations without relying solely on static thresholds. With forecasting, for example, you learn from the historical performance of a time series and predict current and future values. Instead of tuning static thresholds, you can alert when a metric is out of bounds. Forecasting also captures daily and weekly seasonality, which helps with peak vs. off-peak hours, capacity planning, and autoscaling. For example, many services have cyclical trends in traffic (e.g., lower at night or on the weekends, and higher during the weekday). Forecasting can help you alert when traffic is abnormally high or low based on those trends, rather than against a static threshold. With outlier detection, you monitor a group of similar services or instances and identify when one member isn't performing like the rest - the memory leak, the noisy neighbor, the replica that fell behind. For example, you can detect when one or more pods' memory keeps climbing while the rest of the deployment is flat. Observability, simplified: natural-language searches and automated investigations. Querying telemetry can be difficult. You need to understand both your data and the query language before you can start making sense of your metrics, logs, traces, and profiles. Beyond Watchers, Grafana Assistant can take on some of that complexity for you. Here are two examples, whether you want help exploring your data or an agent to investigate an issue on your behalf. Explore your data in natural language: Grafana Assistant. Grafana Assistant lets you interact with your data through natural language. You can ask it to write a query, explain a panel, build a dashboard, or walk you through what's on the page. Because Assistant understands the context of your specific environment, you can use it for broad questions or more specific, focused exploration without worrying about what the underlying tool or query language is. Just give the intent and Assistant does the rest. Use Assistant from the sidebar while you stay on a dashboard, or open Workspace when you want a full-page view with chat, context, and visualizations side by side. You can try Assistant in the public Grafana Play sandbox at play.grafana.com - no install, no account, and no setup required. When you're ready to use it on your own stack, enable it in Grafana Cloud. Assistant is also accessible via self-managed Grafana environments, starting with Grafana 13. Automated root cause analysis: Assistant Investigations. Once you've identified an anomaly, the next question is why it happened. You can use Grafana Assistant to investigate via chat, but what if that analysis could already be underway by the time the pager rings? That's where Assistant Investigations comes in, helping you automatically discover incidents and find root causes faster. Investigations runs longer, prompt-driven analyses across metrics, logs, traces, and profiles within your Grafana Cloud stack. It builds and tests hypotheses, then produces a structured report with findings, evidence, and recommended next steps for incident response. Assistant Investigations lives in Assistant Workspaces. You can start from chat, and share the investigation with your team. You can also configure IRM and alert webhooks so investigations launch automatically when an alert group is created or an incident is updated. Go further with AI in Grafana Cloud. Forecasting, outlier detection, Watchers, and Assistant are just a few of the ways Grafana is using AI to make observability easier. Grafana Cloud includes a growing set of AI capabilities that can help you automate recurring work, observe your own AI applications, and bring Assistant into more of the tools you already use. Here are a few more ways to put AI to work: * Automate recurring work with Automations: Save Assistant prompts and run them on demand or on a schedule for workflows like daily digests and recurring checks. * Use Assistant where you already work: Bring Assistant into the tools and workflows you already use everyday, including Slack, Microsoft Teams, an API, or the CLI. * Connect AI agents to Grafana with gcx and MCP server: Give your AI agents structured access to your Grafana instance and let them interact with your telemetry data. * Observe your own AI agents with Agent Observability: Monitor the AI agents you run in production, including their performance, cost, and quality. Grafana Assistant is the easiest way to get started with metrics, logs, traces, dashboards, and more in Grafana Cloud. Grafana has a generous forever-free tier and plans for every use case. Sign up for free now!
Grafana Labs has crossed 10,000 customers worldwide and surpassed $600 million in annual recurring revenue. The open observability platform provider attributes growth to increased AI adoption, with more than 18,000 organisations now using Grafana Assistant, its AI agent for exploring observability data. The company reports that 57% of organisations are implementing LLM observability. Average products used per contracted customer nearly doubled from 2.3 to 4.4 over two years, with 64.9% now using four or more products. Monthly active Grafana Cloud users grew from roughly 127,000 to over 251,000 over the past two years. The platform's Adaptive Telemetry suite has optimised away 28.5 billion metric series and reduced log volumes by 26PB, cutting telemetry costs by 30 to 50% on average. Grafana Labs was named a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms for the third consecutive year.
Grafana Labs has surpassed $600 million in annual recurring revenue and 10,000 customers globally, up at least 50% from $400 million in September. CEO Raj Dutt attributed the growth largely to AI, as companies move from experimentation to deployment. The observability company helps businesses monitor software systems. AI has increased demand by generating more data and issues to track, such as agents recommending competitors or creating security risks. Grafana launched six AI features in July. Its Agent Observability product monitors AI agents, whilst over 18,000 organisations use Grafana Assistant for system troubleshooting. Dutt said Grafana Assistant is the company's fastest-growing product historically. The company also offers Adaptive Telemetry tools to reduce monitoring costs, which Dutt said cost Grafana nearly $100 million in potential revenue.
Visual playback of the user journey: introducing Session Replay in Grafana Cloud Frontend Observability. 2026-08-17 - 7 min Grafana Cloud Frontend Observability helps engineering teams quantify the end user experience by bringing metrics, logs, traces, and user session context to client-side web applications. Teams can monitor application health and performance over time, triage errors, and correlate frontend signals with backend telemetry to investigate issues across the stack. Yet some of the hardest frontend problems remain difficult to diagnose. A support ticket might report that a checkout button did nothing, a form unexpectedly reset, or a workflow broke only in one browser or on one device. Metrics can reveal a performance regression, logs can capture an error, and traces can expose a slow request, but no single signal shows what the interface actually looked like to the user. This is exactly why Grafana built Session Replay, an add-on feature in Frontend Observability that provides a visual reconstruction of a user's journey, connected to the telemetry Grafana Cloud already collects. It helps engineering teams move from a reported problem to seeing what happened and knowing exactly where to investigate next.
Grafana Labs adds adaptive profiles to observability suite. Wed, 5th Aug 2026 (Today) Grafana Labs has made Adaptive Profiles generally available in Grafana Cloud, completing its Adaptive Telemetry suite across metrics, logs, traces and profiles. The product adjusts profiling detail and collection frequency based on workload behaviour. That lets engineering teams capture more detailed performance data during anomalies while keeping routine collection at a lower-cost baseline. Grafana Labs is framing the launch around rising observability costs as companies collect growing volumes of operational data from software systems. The pressure is increasing as businesses adopt AI agents and applications built with large language models, which generate additional telemetry from model calls, tool use and automated workflows. According to Grafana Labs' 2026 Observability Survey, 57% of organisations are already implementing LLM observability in some form, while 65% say cost is the main criterion when choosing observability tools. Adaptive Telemetry is intended to address that pressure by analysing how telemetry data is used and recommending what should be kept, aggregated or dropped. With the addition of Adaptive Profiles, the approach now spans the four main observability signals used by software teams. Steven Dungan, staff product manager at Grafana Labs, set out the company's position on the economics of observability. "The fundamental problem with observability economics today is that cost scales with ingestion, not insight," Dungan said. "Adaptive Telemetry inverts that model. Every signal - metrics, logs, traces and profiles - now has an intelligent layer that learns how data is used in practice and then optimizes automatically. With Adaptive Profiles reaching GA, we've closed the loop on the full stack. Teams get more signal, less noise and lower bills, and they don't have to sacrifice one for another." Profiles added Continuous profiling shows how applications consume CPU, memory and other resources in production systems. But broad deployment has often been limited by cost, particularly when profiling runs at high resolution across large infrastructure estates. Adaptive Profiles addresses that by varying data collection automatically. Under normal conditions, it gathers profiling data at a lower level. When a performance problem or anomaly appears, it increases the resolution so engineers have more information to investigate. Grafana Labs argues this makes wider use of profiling more financially practical for teams that have struggled to justify fleet-wide deployment. Upland Software said cost control had been a key concern. "Adaptive Profiles ensures that we can leverage Cloud Profiles without worrying about cost overruns," said Michael Beltz, vice president of cloud operations at Upland Software. "The ability to see where code is slowing down, memory is being allocated, and where improvements are needed [with Cloud Profiles] is necessary for us to reduce infrastructure resources and improve the user experience." Existing results Adaptive Profiles joins three other tools in the suite that are already generally available: Adaptive Metrics, Adaptive Logs and Adaptive Traces. According to Grafana Labs, those products have delivered measurable reductions in data volumes and spending among Grafana Cloud users. Adaptive Metrics is the most widely deployed of the four, the company said. It has removed 28.5 billion active series and delivered an average 35% reduction in metrics costs. Mux was cited as one customer that cut its metrics volume by 60% and extended retention from 14 days to 13 months. "Adaptive Metrics is an amazing feature. It not only saves us hundreds of thousands of dollars a year, but it's also a forcing function for us to look closely at our metrics to find additional opportunities for time series reduction and cardinality improvements," said Kyle Weaver, staff software engineer at Mux. Adaptive Logs examines log usage to identify high-volume patterns that are rarely used and can be removed. Grafana Labs said the feature has eliminated 26 petabytes of log volume across its cloud platform. TeleTracking, an early adopter of the tool, has seen a 50% reduction in log volumes, according to Grafana Labs. "Adaptive Logs helps reduce noise, making it easier to spot valuable logs and ultimately saves us costs," said Andrew Qu, software engineer II at TeleTracking. Adaptive Traces, which became generally available in late 2025, uses tail sampling to keep traces linked to errors, latency and other notable events while filtering lower-value, repetitive spans. Grafana Labs said this has reduced trace data volume by an average of 82% across Grafana Cloud. Auditboard said the tool changed the trade-off between visibility and cost. "Before Adaptive Traces, we had two bad options: send everything and blow our budget, or send so little we couldn't get meaningful insight," said Geoff Schultz, manager of infrastructure engineering at Auditboard. "Now tracing is actually usable, we can dial sampling up or down as needed, keep costs in check and still give teams the visibility they need." Across customers using multiple parts of the Adaptive Telemetry suite, Grafana Labs said total telemetry costs have fallen by 30% to 50% on average, with savings redirected into broader observability coverage.
Find jobs on Simplify and start your career today
Industries
Data & Analytics
Enterprise Software
Company Size
1,001-5,000
Company Stage
Series D
Total Funding
$805.2M
Headquarters
New York City, New York
Founded
2014
Find jobs on Simplify and start your career today