Full-Time
Open-source dashboards and cloud observability platform
€94k - €112.8k/yr
Remote in Spain
Remote
Remote role; applicants located in Spain, Sweden, United Kingdom, Ireland, or Germany.
See people who can refer or advise you
Grafana Labs builds observability and monitoring tools for cloud infrastructure and applications. Its flagship Grafana dashboard lets users visualize data from many sources in real time and set up alerts, with additional options like Grafana Enterprise for large deployments and Grafana Cloud as a managed service. The core open-source platform is complemented by commercial features and services that provide security, scalability, and dedicated support, appealing to both individual developers and large organizations. The goal is to help businesses keep digital services reliable and efficient by delivering scalable, real-time visibility into software and infrastructure.
Company Size
1,001-5,000
Company Stage
Series D
Total Funding
$805.2M
Headquarters
New York City, New York
Founded
2014
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
30 days of paid vacation each year on top of national holidays, parental leave, & sick leave
Health coverage
4% contribution match on our 401(k)
$1,500 learning and development stipend
Udemy subscription
Complimentary subscription to Headspace
Discounts on a wide variety of services, including entertainment, food, and fitness.
Remote Work Option
Global Employee Assistance Program
Visual playback of the user journey: introducing Session Replay in Grafana Cloud Frontend Observability. 2026-08-17 - 7 min Grafana Cloud Frontend Observability helps engineering teams quantify the end user experience by bringing metrics, logs, traces, and user session context to client-side web applications. Teams can monitor application health and performance over time, triage errors, and correlate frontend signals with backend telemetry to investigate issues across the stack. Yet some of the hardest frontend problems remain difficult to diagnose. A support ticket might report that a checkout button did nothing, a form unexpectedly reset, or a workflow broke only in one browser or on one device. Metrics can reveal a performance regression, logs can capture an error, and traces can expose a slow request, but no single signal shows what the interface actually looked like to the user. This is exactly why Grafana built Session Replay, an add-on feature in Frontend Observability that provides a visual reconstruction of a user's journey, connected to the telemetry Grafana Cloud already collects. It helps engineering teams move from a reported problem to seeing what happened and knowing exactly where to investigate next.
Grafana Labs adds adaptive profiles to observability suite. Wed, 5th Aug 2026 (Today) Grafana Labs has made Adaptive Profiles generally available in Grafana Cloud, completing its Adaptive Telemetry suite across metrics, logs, traces and profiles. The product adjusts profiling detail and collection frequency based on workload behaviour. That lets engineering teams capture more detailed performance data during anomalies while keeping routine collection at a lower-cost baseline. Grafana Labs is framing the launch around rising observability costs as companies collect growing volumes of operational data from software systems. The pressure is increasing as businesses adopt AI agents and applications built with large language models, which generate additional telemetry from model calls, tool use and automated workflows. According to Grafana Labs' 2026 Observability Survey, 57% of organisations are already implementing LLM observability in some form, while 65% say cost is the main criterion when choosing observability tools. Adaptive Telemetry is intended to address that pressure by analysing how telemetry data is used and recommending what should be kept, aggregated or dropped. With the addition of Adaptive Profiles, the approach now spans the four main observability signals used by software teams. Steven Dungan, staff product manager at Grafana Labs, set out the company's position on the economics of observability. "The fundamental problem with observability economics today is that cost scales with ingestion, not insight," Dungan said. "Adaptive Telemetry inverts that model. Every signal - metrics, logs, traces and profiles - now has an intelligent layer that learns how data is used in practice and then optimizes automatically. With Adaptive Profiles reaching GA, we've closed the loop on the full stack. Teams get more signal, less noise and lower bills, and they don't have to sacrifice one for another." Profiles added Continuous profiling shows how applications consume CPU, memory and other resources in production systems. But broad deployment has often been limited by cost, particularly when profiling runs at high resolution across large infrastructure estates. Adaptive Profiles addresses that by varying data collection automatically. Under normal conditions, it gathers profiling data at a lower level. When a performance problem or anomaly appears, it increases the resolution so engineers have more information to investigate. Grafana Labs argues this makes wider use of profiling more financially practical for teams that have struggled to justify fleet-wide deployment. Upland Software said cost control had been a key concern. "Adaptive Profiles ensures that we can leverage Cloud Profiles without worrying about cost overruns," said Michael Beltz, vice president of cloud operations at Upland Software. "The ability to see where code is slowing down, memory is being allocated, and where improvements are needed [with Cloud Profiles] is necessary for us to reduce infrastructure resources and improve the user experience." Existing results Adaptive Profiles joins three other tools in the suite that are already generally available: Adaptive Metrics, Adaptive Logs and Adaptive Traces. According to Grafana Labs, those products have delivered measurable reductions in data volumes and spending among Grafana Cloud users. Adaptive Metrics is the most widely deployed of the four, the company said. It has removed 28.5 billion active series and delivered an average 35% reduction in metrics costs. Mux was cited as one customer that cut its metrics volume by 60% and extended retention from 14 days to 13 months. "Adaptive Metrics is an amazing feature. It not only saves us hundreds of thousands of dollars a year, but it's also a forcing function for us to look closely at our metrics to find additional opportunities for time series reduction and cardinality improvements," said Kyle Weaver, staff software engineer at Mux. Adaptive Logs examines log usage to identify high-volume patterns that are rarely used and can be removed. Grafana Labs said the feature has eliminated 26 petabytes of log volume across its cloud platform. TeleTracking, an early adopter of the tool, has seen a 50% reduction in log volumes, according to Grafana Labs. "Adaptive Logs helps reduce noise, making it easier to spot valuable logs and ultimately saves us costs," said Andrew Qu, software engineer II at TeleTracking. Adaptive Traces, which became generally available in late 2025, uses tail sampling to keep traces linked to errors, latency and other notable events while filtering lower-value, repetitive spans. Grafana Labs said this has reduced trace data volume by an average of 82% across Grafana Cloud. Auditboard said the tool changed the trade-off between visibility and cost. "Before Adaptive Traces, we had two bad options: send everything and blow our budget, or send so little we couldn't get meaningful insight," said Geoff Schultz, manager of infrastructure engineering at Auditboard. "Now tracing is actually usable, we can dial sampling up or down as needed, keep costs in check and still give teams the visibility they need." Across customers using multiple parts of the Adaptive Telemetry suite, Grafana Labs said total telemetry costs have fallen by 30% to 50% on average, with savings redirected into broader observability coverage.
Grafana Labs announced the general availability of Adaptive Profiles in Grafana Cloud, completing its Adaptive Telemetry suite across all four observability signals: metrics, logs, traces, and profiles. The suite automatically identifies and retains high-value data whilst filtering out unnecessary information without manual engineering work. According to Grafana Labs' 2026 Observability Survey, 57% of organisations are implementing LLM observability, with 65% citing cost as the top criterion for selecting tools. Adaptive Profiles dynamically adjusts data collection based on workload behaviour, collecting at baseline during normal operations and increasing resolution when anomalies arise. The complete suite has delivered measurable results: Adaptive Metrics eliminated 28.5 billion active series with an average 35% cost reduction, Adaptive Logs removed 26 petabytes of volume, and Adaptive Traces reduced data volume by 82% on average. Organisations using multiple components see 30–50% reductions in total telemetry costs.
Grafana Labs gives every observability signal an intelligent optimization layer, completing Adaptive Telemetry suite with Adaptive Profiles GA. Published: 4 Aug 2026 Grafana Labs newsroom press releases Grafana Labs gives every observability signal an intelligent optimization layer, completing Adaptive Telemetry suite with Adaptive Profiles GA. With Grafana Labs' Adaptive Telemetry suite, customers including Mux, SailPoint, TeleTracking, and Auditboard have cut telemetry spend by 30-50% on average, redirecting savings into deeper observability coverage NEW YORK - August 4, 2026 - Grafana Labs, the company behind the open observability cloud, today announced the general availability of Adaptive Profiles in Grafana Cloud, completing the Adaptive Telemetry suite to span all four telemetry signals: metrics, logs, traces, and profiles. With Adaptive Profiles now GA, every layer of an organization's observability stack can automatically identify and retain high-value data while filtering out what doesn't matter, without manual tuning or engineering toil. Telemetry volumes are growing faster than the insight they generate, and AI is accelerating the problem. As organizations deploy AI agents and LLM-powered applications, every model call, tool invocation, and agentic workflow generates new telemetry that needs to be observed, traced, and profiled. According to Grafana Labs' 2026 Observability Survey, 57% of organizations are already implementing LLM observability in some capacity, and 65% cite cost as the top criteria for selecting observability tools. AI gives teams more to watch, which means the cost of watching everything is rising sharply. Most organizations are already collecting far more data than they ever query, alert on, or act upon - AI-generated telemetry is making that gap wider, faster. Grafana Labs built the Adaptive Telemetry suite to address this directly: not by asking teams to manually cull their data, but by continuously analyzing how telemetry is actually used and surfacing precise recommendations for what to keep, aggregate, or drop. "The fundamental problem with observability economics today is that cost scales with ingestion, not insight," said Steven Dungan, Staff Product Manager at Grafana Labs. "Adaptive Telemetry inverts that model. Every signal: metrics, logs, traces, and profiles, now has an intelligent layer that learns how data is used in practice, and then optimizes automatically. With Adaptive Profiles reaching GA, we've closed the loop on the full stack. Teams get more signal, less noise, and lower bills and they don't have to sacrifice one for another." Adaptive Profiles: Continuous profiling at scale, without runaway costs. Continuous profiling gives engineers deep insight into how applications consume CPU, memory, and other resources in production - but broad deployment across infrastructure has historically been cost-prohibitive. Adaptive Profiles changes this by dynamically adjusting the detail and frequency of data collection based on workload behavior. During normal operations, it collects at a cost-effective baseline. When anomalies or performance issues arise, it automatically increases resolution to ensure engineers have the data they need to investigate. For engineering teams that have struggled to justify fleet-wide profiling, Adaptive Profiles makes the economics work, delivering richer performance data where it matters, without requiring a fixed high-cost collection rate across every service. "Adaptive Profiles ensures that we can leverage Cloud Profiles without worrying about cost overruns," said Michael Beltz, VP, Cloud Operations at Upland Software. "The ability to see where code is slowing down, memory is being allocated, and where improvements are needed [with Cloud Profiles] is necessary for us to reduce infrastructure resources and improve the user experience." The complete suite: savings across every signal. Adaptive Profiles joins three generally available capabilities that have already delivered significant, measurable results across the Grafana Cloud customer base. Adaptive Metrics is now the most widely deployed component of the suite and has helped customers eliminate 28.5 billion active series, delivering an average 35% reduction in metrics costs. One customer, Mux, was able to cut its metrics volume by 60% and extended retention from 14 days to 13 months. "Adaptive Metrics is an amazing feature. It not only saves us hundreds of thousands of dollars a year, but it's also a forcing function for us to look closely at our metrics to find additional opportunities for time series reduction and cardinality improvements," said Kyle Weaver, Staff Software Engineer at Mux. Adaptive Logs applies the same usage-based analysis to log data, identifying high-volume, low-value log patterns and generating recommendations for what can be safely dropped. Across Grafana Cloud, Adaptive Logs has eliminated 26 petabytes of log volume - data that teams were paying to store but never using. TeleTracking, an early adopter, has seen a 50% reduction in log volumes. "Adaptive Logs helps reduce noise, making it easier to spot valuable logs and ultimately saves us costs," said Andrew Qu, Software Engineer II at TeleTracking. Adaptive Traces, which reached general availability in October 2025, uses tail sampling to ensure that traces with errors, high latency, or other signals of interest are always retained, while noise from healthy, repetitive spans is filtered out. Across Grafana Cloud, Adaptive Traces has reduced trace data volume by an average of 82%. "Before Adaptive Traces, we had two bad options: send everything and blow our budget, or send so little we couldn't get meaningful insight," said Geoff Schultz, Manager, Infrastructure Engineering at Auditboard. "Now tracing is actually usable, we can dial sampling up or down as needed, keep costs in check, and still give teams the visibility they need." Across the full Adaptive Telemetry suite, organizations using multiple components see 30-50% reductions in total telemetry costs on average. About Grafana Labs. Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, its fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions. Today, more than 35 million users and 7,000+ customers - including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce - trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. Grafana Labs Inc. is a 100% remote company with 1,400+ team members across 40+ countries, and Grafana Labs Inc. is backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital. Learn more at grafana.com and follow Grafana Labs Inc. on LinkedIn and X. Press Contact: [email protected]
Grafana Labs launches six AI tools to power agentic operations. Grafana Labs, the open observability cloud provider, announced the general availability of six artificial intelligence capabilities launched during its inaugural AI Week. The releases expand Grafana Assistant into a comprehensive agentic operations layer designed to detect, investigate, and remediate production issues at the accelerated speed of modern AI development. The recently launched set includes Grafana Assistant Investigations, Grafana Assistant Workspace, Grafana Assistant Automations, the Grafana Cloud Model Context Protocol (MCP) Server, gcx, and Grafana Agent Observability. As a result, these technologies help transform the observability approach to a lifecycle one by implementing it at different stages, including after the deployment of the product. "We used to treat observability as something you bolt on just before code reaches production," said Mat Ryer, Senior Director of AI, Grafana Labs. "That's changing. Now, Grafana Assistant can review your plans before you've written a line of code, add the instrumentation for you once you have, watch features as they land in production, and stay with you as you scale while dealing with the inevitable incidents that follow." Bridging the observability gap across the software lifecycle. Traditional observability has historically operated after code reaches production environments - requiring engineers to manually instrument services, construct dashboards, and monitor alerts. However, as AI coding tools dramatically increase deployment velocity, operational practices must evolve to prevent production bottlenecks. According to Grafana Labs' 2026 Observability Survey, 92% of practitioners report seeing real value in AI detecting system anomalies, yet only 57% are currently implementing observability for their internal AI systems. Grafana's new capabilities address this adoption gap by automating critical operational workflows: Plan and Code Instrumentation: Using gcx - Grafana's agentic CLI - and the Grafana Cloud MCP server, engineers can manage dashboards, data sources, and alert rules as code. The Assistant can automatically review architectural plans, open pull requests with necessary instrumentation, and verify telemetry arrival. Automated Investigation and Remediation: When production anomalies occur, Grafana Assistant Investigations automatically generates diagnostic hypotheses, analyzes underlying telemetry data, and presents validated conclusions to engineers. AI System Observability: Through Grafana Agent Observability, teams can monitor OpenTelemetry-native metrics across custom AI applications, tracking key performance signals such as latency, token expenditure, conversation histories, and model behavior. "AI should allow you to move at 10x velocity, not produce incidents at 10x the rate," Ryer said. "That's why we're releasing all these AI capabilities in a single week. Engineers shouldn't have to wait for their tools to catch up to how quickly they're already moving."