Full-Time

Staff Backend Engineer

Second Horizon

Updated on 9/4/2026

Grafana Labs

Grafana Labs

1,001-5,000 employees

Open-source dashboards and cloud observability platform

Compensation Overview

$175k - $210k/yr

+ Equity + Bonus

Remote in USA

Remote

Applicants must be in the U.S. or Canada; Quebec residents are not eligible. In-person onboarding is required.

Category
Software Engineering (1)
Required Skills
LLM
Power BI
Microsoft Azure
Data Visualization
Data Engineering
Tableau
AWS
Observability
Databricks
Looker
Data Analysis
Snowflake
Google Cloud Platform

Get referred to Grafana Labs

See people who can refer or advise you

Requirements
  • Experience with large language models, prompt engineering, and building applications powered by generative artificial intelligence.
  • A proven track record of delivering software that reached production and is actively used by users.
  • Exposure to cloud-native environments such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure.
  • Experience using observability tools to understand and troubleshoot system behavior.
Responsibilities
  • Design, implement, test, and operate backend services for context ingestion, context indexing, retrieval orchestration, application programming interface access, source configuration, and system administration.
  • Define and build the architecture for a multi-tenant software-as-a-service foundation, including tenant isolation, usage tracking, quotas, audit logs, background jobs, and reliable service boundaries.
  • Build application programming interfaces and service interfaces that enable artificial intelligence agents, Model Context Protocol tools, command-line interfaces, and internal applications to retrieve relevant context, provenance, confidence signals, and warnings.
  • Partner across product and infrastructure to make practical tradeoffs between rapid experimentation and long-term reliability as the project moves from prototype to production.
  • Instrument services with metrics, logs, traces, alerts, and dashboards; use observability tools to understand system behavior and improve reliability.
  • Help shape the architecture, service boundaries, storage choices, application programming interface contracts, deployment patterns, and engineering practices for a new product area.
  • Communicate effectively and contribute across teams in a dynamic, collaborative environment.
  • Take full ownership of the artificial intelligence solutions developed, ensuring they are scalable, maintainable, and aligned with real user workflows.
Desired Qualifications
  • Experience building or working with agent frameworks or multi-agent workflows.
  • Experience as a data analyst or working with data platforms such as Looker, Tableau, Power BI, Snowflake, or Databricks.
  • Experience building tools for data engineering.

Grafana Labs builds observability and monitoring tools for cloud infrastructure and applications. Its flagship Grafana dashboard lets users visualize data from many sources in real time and set up alerts, with additional options like Grafana Enterprise for large deployments and Grafana Cloud as a managed service. The core open-source platform is complemented by commercial features and services that provide security, scalability, and dedicated support, appealing to both individual developers and large organizations. The goal is to help businesses keep digital services reliable and efficient by delivering scalable, real-time visibility into software and infrastructure.

Company Size

1,001-5,000

Company Stage

Series D

Total Funding

$805.2M

Headquarters

New York City, New York

Founded

2014

Get referred to Grafana Labs

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • August 26, 2026: Grafana reported $600M ARR, up from $400M last year.
  • July 2026 AI Week launched six GA products, deepening workflow lock-in.
  • Adaptive Profiles GA on August 4 completed full-stack cost optimization across telemetry.

What critics are saying

  • CVE-2026-19516 hit Grafana MCP on August 11, exposing internal networks to SSRF.
  • September 2, 2026 advisories added SAML replay and auth-proxy session takeover bugs.
  • Repeated MCP vulnerabilities can freeze enterprise rollouts of Assistant and stall Grafana Cloud expansion.

What makes Grafana Labs unique

  • Grafana’s open-source footprint spans 10,000 customers and 35 million users by August 2026.
  • Adaptive Telemetry cuts customer telemetry costs 30-50% across metrics, logs, traces, profiles.
  • AI Week’s Agentic Operations stack adds Investigations, MCP, gcx, and Agent Observability.

Help us improve and share your feedback! Did you find this helpful?

Benefits

30 days of paid vacation each year on top of national holidays, parental leave, & sick leave

Health coverage

4% contribution match on our 401(k)

$1,500 learning and development stipend

Udemy subscription

Complimentary subscription to Headspace

Discounts on a wide variety of services, including entertainment, food, and fitness.

Remote Work Option

Global Employee Assistance Program

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

1%

2 year growth

0%
Content+Technology
Sep 3rd, 2026
Imagine Communications unveils latest control portfolio advances.

Imagine Communications unveils latest control portfolio advances. 04/09/2026 At IBC2026 (stand 1.B73), Imagine Communications is showcasing the latest advances in its broad portfolio of industry-proven control tools that enable users to seamlessly manage, monitor and orchestrate IP, cloud and hybrid operations and build a practical path to AI-enabled workflows and future Dynamic Media Facilities (DMF). New at IBC2026 is Magellan PanelFlex, the next generation of operator-focused, freeform control surfaces for Imagine's widely deployed Magellan Control System. Fully context-based, Magellan PanelFlex adapts controls to dynamic production environments, ensuring operators always have access to the tools they need, exactly when they need them. Magellan PanelFlex enables freeform naming of signals at every operating position, for every user, with no limits. The Magellan Control System delivers IP/SDI routing control and real-time monitoring and an operator-centric work surface that makes the IP technology underneath feel and act just like a traditional SDI router. Magellan PanelFlex adds a full ecosystem of advanced capabilities, including Magellan PanelFlex-Canvas - a powerful panel design environment that allows engineering users to sculpt the exact flow of control and information for each operating position. The Magellan PanelFlex real-time engine drives these canvases to operators through HTML5 web panels, dynamic integration onto multiviewer surfaces and familiar physical panels. "Our mission with Magellan PanelFlex is to put operators in full control and simplify the live event production process," said Adeel Ahmed, Senior Product Manager, Magellan Control System, at Imagine Communications. "From reassigning name sets to reconfiguring buttons on individual control surfaces, everything is just a click away - so the show goes faster and smoother." In keeping with Imagine's commitment to open interfaces and protocols, Magellan PanelFlex-AI organises and curates all key system information into an agentic-enabling Model Context Protocol (MCP) interface, allowing AI systems to quickly follow and act upon status and information from the live system. New capabilities across the control portfolio. Imagine is introducing these MCP interfaces across its portfolio of workflow-optimised control technologies, which together form the software layer that connects operators to content and infrastructure and provides a practical foundation for the next generation of media operations. For systems deployment and management, Imagine's Aviator Orchestrator streamlines activities across on-premises, cloud and hybrid environments, creating a practical bridge to the DMF vision. Augmenting Imagine control capabilities, the software-defined Prismon multiviewer provides centralised monitoring, fast diagnostics and unlimited scalability - helping operators maintain confidence in modern media workflows by enabling easy management of all displays from a single intuitive interface. Underpinning the full portfolio, Imagine has extended Prometheus monitoring endpoints across its product suite, enabling users to integrate with IT-standard observability platforms such as Grafana. Imagine's Professional Services team works directly with customers to build tailored dashboards for at-a-glance visibility across the entire facility. "Our customers are managing increasingly complex facilities with shrinking operational teams, and our control portfolio is designed for exactly that challenge," said John Mailhot, Senior Vice President, Product Management, at Imagine Communications. "We build to open interfaces and standard protocols, so we can provide customers with systemic monitoring and unified visibility leveraging available IT tools, rather than asking them to adopt something proprietary. At IBC, we'll be demonstrating how those same open interfaces provide a foundation for AI-enabled workflows, making trusted systems smarter, faster and easier to run." Imagine will demonstrate Magellan PanelFlex alongside its full control portfolio and AI-enabled technologies at its IBC2026 stand (1.B73).

SC Media
Sep 3rd, 2026
Grafana fixes critical SSRF flaw affecting Grafana MCP servers.

Grafana fixes critical SSRF flaw affecting Grafana MCP servers. September 3, 2026 (Credit: frank - stock.adobe.com) Two weaknesses in the Grafana MCP server implementation, including a critical server-side request forgery (SSRF) flaw, were patched by Grafana last month, according to a report by Pillar Security published Wednesday. The SSRF vulnerability, tracked as CVE-2026-19516 with a CVSS score of 9.1, could have allowed an MCP caller to leverage the Grafana MCP server's network position to make requests to internal addresses that would not have otherwise been reachable. Pillar Security, which discovered and reported the flaw, also reported an additional weakness that allowed an unauthenticated caller to generate their own session ID and use it to call tools via the Grafana MCP. This was due to the MCP server validating the session ID format rather than checking for a valid credential, allowing "session-shaped IDs" that the server never issued to be accepted, Pillar researchers explained. These unauthenticated tool calls could be paired with the SSRF vulnerability to form a "killchain," Pillar said, potentially allowing an unauthorized attacker to retrieve sensitive information from the internal network hosting the MCP server. Grafana added optional bearer-token protection in Grafana MCP v 1.1.0 to address the session validation issue, although the company considered this "a security improvement rather than a vulnerability," Pillar said. Related reading: CVE-2026-19516, also fixed in v1.1.0, allowed a caller to supply an "X-Grafana-URL" request header when calling the "grafana_api_request" tool, which enabled them to control the HTTP method, path and body of the outbound request. An attacker could receive the output of requests to internal services reachable by the Grafana MCP server, leveraging the server's network position to reach otherwise inaccessible destinations. Pillar demonstrated in tests how this could be exploited in certain setups to retrieve cloud credentials from the AWS Instance Metadata Service (IMDSv2), first using a PUT request and TTL header to obtain the metadata session token and then issuing a second request leveraging that token to retrieve the credentials. "This does not mean that every Grafana MCP deployment could reach cloud metadata. However, it shows the relevant security property: the server can become a readable, method-capable proxy from its own network location," Pillar AI Security Researcher Ariel Fogel wrote in the report. Grafana previously patched a vulnerability tracked as CVE-2026-15583, preventing the Grafana service-account token from being exfiltrated to external websites via a crafted X-Grafana-URL request header. Pillar confirmed that the flaw they discovered did not allow the caller to obtain the service-account token itself but noted that the MCP server's "authentication posture" could still be exploited as a proxy to sensitive internal data. Pillar concluded that organizations should view MCP servers as "identity brokers," giving human and AI callers the ability to perform actions under the server's privileges and network access. They recommend that tools capable of issuing HTTP requests should have an explicit destination policy and that remote MCP deployments should require authentication. "The caller provides the instruction and the server provides the reach. Security depends on keeping those two things connected," Pillar concluded.

Grafana Labs
Aug 27th, 2026
Identify anomalies, outlier detection, forecasting: How Grafana Cloud uses AI to make observability easier.

Identify anomalies, outlier detection, forecasting: How Grafana Cloud uses AI to make observability easier. 2026-08-27 - 6 min Note: This blog originally published in July 2024 and was updated in August 2026 to reflect the latest AI solutions and workflows in Grafana Cloud. Tech stacks keep getting more complicated, which means they're also more difficult to monitor. At Grafana Labs, Grafana is building AI solutions to help you understand these complex systems with less toil, so you can get to answers faster when something looks off. A big part of that is helping you catch anomalies earlier and pinpoint root causes, without needing to be an expert in every query language. In this post, Grafana'll explore a few of the AI-powered capabilities in Grafana Cloud designed to do just that, from forecasting and outlier detection to easier data exploration and incident response with Grafana Assistant, Watchers, and Investigations. Find anomalies: Agentic alerting, forecasting, and outlier detection. There are lots of ways AI can help you spot potential issues earlier in Grafana Cloud. Here are a few capabilities that help you identify when your system deviates from what's expected, whether that's against your own definition of "normal," against historical patterns, or compared with similar services and instances. Agentic alerting for your systems: Assistant Watchers. What if you know what you want to keep an eye on, but you don't know exactly which metrics or thresholds to alert on? That's why Grafana built Assistant Watchers. A capability of Grafana Assistant, the AI-powered agent in Grafana Cloud (more on that below), Watchers are always-on agents that check a scoped slice of your telemetry on a schedule. You describe what "normal" looks like in your own words - the service, the failure mode, the noise that's expected during a deployment. Assistant Watchers calibrates concrete checks against your metrics and logs, then raises concern when a run finds something notable. For example, let's say you deployed a new service and have some telemetry, but you aren't sure what to alert on yet. You can use a watcher to monitor the service and tune what "normal" looks like as you go. A watcher can stay quiet when signals look healthy, notify via Slack or a webhook when it flags a warning, and launch Assistant Investigations (again, more on that below) for critical findings. Use Watchers for recurring service health checks, noisy-but-important failure modes, and follow-up monitoring after an incident or deployment. Forecasting and outlier detection. When you already know which metrics matter, forecasting and outlier detection features in Grafana Cloud give you more targeted ways to identify unusual behavior. These tried-and-true approaches learn what to expect from your metrics over time or across groups of similar services, so you can detect deviations without relying solely on static thresholds. With forecasting, for example, you learn from the historical performance of a time series and predict current and future values. Instead of tuning static thresholds, you can alert when a metric is out of bounds. Forecasting also captures daily and weekly seasonality, which helps with peak vs. off-peak hours, capacity planning, and autoscaling. For example, many services have cyclical trends in traffic (e.g., lower at night or on the weekends, and higher during the weekday). Forecasting can help you alert when traffic is abnormally high or low based on those trends, rather than against a static threshold. With outlier detection, you monitor a group of similar services or instances and identify when one member isn't performing like the rest - the memory leak, the noisy neighbor, the replica that fell behind. For example, you can detect when one or more pods' memory keeps climbing while the rest of the deployment is flat. Observability, simplified: natural-language searches and automated investigations. Querying telemetry can be difficult. You need to understand both your data and the query language before you can start making sense of your metrics, logs, traces, and profiles. Beyond Watchers, Grafana Assistant can take on some of that complexity for you. Here are two examples, whether you want help exploring your data or an agent to investigate an issue on your behalf. Explore your data in natural language: Grafana Assistant. Grafana Assistant lets you interact with your data through natural language. You can ask it to write a query, explain a panel, build a dashboard, or walk you through what's on the page. Because Assistant understands the context of your specific environment, you can use it for broad questions or more specific, focused exploration without worrying about what the underlying tool or query language is. Just give the intent and Assistant does the rest. Use Assistant from the sidebar while you stay on a dashboard, or open Workspace when you want a full-page view with chat, context, and visualizations side by side. You can try Assistant in the public Grafana Play sandbox at play.grafana.com - no install, no account, and no setup required. When you're ready to use it on your own stack, enable it in Grafana Cloud. Assistant is also accessible via self-managed Grafana environments, starting with Grafana 13. Automated root cause analysis: Assistant Investigations. Once you've identified an anomaly, the next question is why it happened. You can use Grafana Assistant to investigate via chat, but what if that analysis could already be underway by the time the pager rings? That's where Assistant Investigations comes in, helping you automatically discover incidents and find root causes faster. Investigations runs longer, prompt-driven analyses across metrics, logs, traces, and profiles within your Grafana Cloud stack. It builds and tests hypotheses, then produces a structured report with findings, evidence, and recommended next steps for incident response. Assistant Investigations lives in Assistant Workspaces. You can start from chat, and share the investigation with your team. You can also configure IRM and alert webhooks so investigations launch automatically when an alert group is created or an incident is updated. Go further with AI in Grafana Cloud. Forecasting, outlier detection, Watchers, and Assistant are just a few of the ways Grafana is using AI to make observability easier. Grafana Cloud includes a growing set of AI capabilities that can help you automate recurring work, observe your own AI applications, and bring Assistant into more of the tools you already use. Here are a few more ways to put AI to work: * Automate recurring work with Automations: Save Assistant prompts and run them on demand or on a schedule for workflows like daily digests and recurring checks. * Use Assistant where you already work: Bring Assistant into the tools and workflows you already use everyday, including Slack, Microsoft Teams, an API, or the CLI. * Connect AI agents to Grafana with gcx and MCP server: Give your AI agents structured access to your Grafana instance and let them interact with your telemetry data. * Observe your own AI agents with Agent Observability: Monitor the AI agents you run in production, including their performance, cost, and quality. Grafana Assistant is the easiest way to get started with metrics, logs, traces, dashboards, and more in Grafana Cloud. Grafana has a generous forever-free tier and plans for every use case. Sign up for free now!

Associated Press
Aug 26th, 2026
Grafana Labs hits 10,000 customers and $600M ARR as AI observability adoption surges

Grafana Labs has crossed 10,000 customers worldwide and surpassed $600 million in annual recurring revenue. The open observability platform provider attributes growth to increased AI adoption, with more than 18,000 organisations now using Grafana Assistant, its AI agent for exploring observability data. The company reports that 57% of organisations are implementing LLM observability. Average products used per contracted customer nearly doubled from 2.3 to 4.4 over two years, with 64.9% now using four or more products. Monthly active Grafana Cloud users grew from roughly 127,000 to over 251,000 over the past two years. The platform's Adaptive Telemetry suite has optimised away 28.5 billion metric series and reduced log volumes by 26PB, cutting telemetry costs by 30 to 50% on average. Grafana Labs was named a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms for the third consecutive year.

Business Insider
Aug 26th, 2026
Grafana hits $600M ARR as companies shift from AI experimentation to deployment

Grafana Labs has surpassed $600 million in annual recurring revenue and 10,000 customers globally, up at least 50% from $400 million in September. CEO Raj Dutt attributed the growth largely to AI, as companies move from experimentation to deployment. The observability company helps businesses monitor software systems. AI has increased demand by generating more data and issues to track, such as agents recommending competitors or creating security risks. Grafana launched six AI features in July. Its Agent Observability product monitors AI agents, whilst over 18,000 organisations use Grafana Assistant for system troubleshooting. Dutt said Grafana Assistant is the company's fastest-growing product historically. The company also offers Adaptive Telemetry tools to reduce monitoring costs, which Dutt said cost Grafana nearly $100 million in potential revenue.