Full-Time

Engineering Manager

Adaptive Telemetry

Posted on 2/13/2026

Grafana Labs

Grafana Labs

1,001-5,000 employees

Open-source dashboards and cloud observability platform

Compensation Overview

CA$186.4k - CA$223.6k/yr

+ Equity + Bonus

Remote in Canada

Remote

Category
Engineering Management (1)
Required Skills
Observability

Get referred to Grafana Labs

Find people who can refer or advise you

Requirements
  • A motivated self-starter with a strong bias toward action and ownership, able to navigate ambiguity and drive work forward in complex problem spaces.
  • Comfortable collaborating across time zones and cultures in a fully remote environment, using clear, empathetic communication to build trust, connection, and psychological safety.
  • Experienced in guiding and supporting other engineers, leading initiatives or small teams, and collaborating effectively to drive projects to successful outcomes.
  • Demonstrated ability to manage complexity, align stakeholders, and elevate team performance to deliver meaningful, high-quality results.
  • Deeply customer-focused, consistently grounding decisions and technical choices in user needs and real-world impact.
  • Solid understanding of distributed systems principles, including scalability, fault tolerance, consistency, and observability.
  • Comfort with AI-assisted tooling. We embrace AI and agentic development so we expect you to be curious and comfortable using AI-powered tools and ideally have practical experience folding them into a team’s workflow.
Responsibilities
  • Manage and develop a team of engineers, providing regular feedback and supporting each person’s growth through career conversations.
  • Collaborate closely with product, design and engineering leadership to define goals that move the Adaptive Telemetry group forward and align with broader business objectives.
  • Provide guidance throughout the project lifecycle from early ideation through planning, execution and post‑launch guiding the team to deliver high‑quality results consistently
  • Foster a psychologically safe environment where engineers can learn, experiment and iterate quickly, encouraging innovation and a culture of continuous improvement.
  • Work across team boundaries by contributing to projects outside your team’s scope and partnering with other functions to solve customer problems.

Grafana Labs builds observability and monitoring tools for cloud infrastructure and applications. Its flagship Grafana dashboard lets users visualize data from many sources in real time and set up alerts, with additional options like Grafana Enterprise for large deployments and Grafana Cloud as a managed service. The core open-source platform is complemented by commercial features and services that provide security, scalability, and dedicated support, appealing to both individual developers and large organizations. The goal is to help businesses keep digital services reliable and efficient by delivering scalable, real-time visibility into software and infrastructure.

Company Size

1,001-5,000

Company Stage

Series D

Total Funding

$805.2M

Headquarters

New York City, New York

Founded

2014

Get referred to Grafana Labs

Find people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • AI firms like 7AI and Zama standardize on Grafana Cloud to unify observability and cut OSS overhead.
  • Managed/SaaS observability adoption rose to 50% in 2026, fueling Grafana Cloud subscription growth.
  • Tempo 3.0’s Kafka-compatible architecture reduces storage costs by storing one trace copy instead of three.

What critics are saying

  • TanStack supply chain breach via Mini Shai-Hulud enabled GitHub token theft and codebase exfiltration in May 2026.
  • CoinbaseCartel data-extortion group targets developer environments using stolen credentials with high impact probability.
  • Datadog’s Q1 2026 LLM Observability and GPU Monitoring capture AI workload growth narrative, threatening Grafana’s market share.

What makes Grafana Labs unique

  • Grafana Labs combines open-source Grafana with fully managed Cloud and Enterprise offerings for unified observability.
  • LGTM Stack (Mimir, Loki, Tempo) delivers scalable metrics, logs, and traces in one composable platform.
  • Grafana Assistant enables natural-language querying and AI-driven root-cause analysis without deep PromQL expertise.

Help us improve and share your feedback! Did you find this helpful?

Benefits

30 days of paid vacation each year on top of national holidays, parental leave, & sick leave

Health coverage

4% contribution match on our 401(k)

$1,500 learning and development stipend

Udemy subscription

Complimentary subscription to Headspace

Discounts on a wide variety of services, including entertainment, food, and fitness.

Remote Work Option

Global Employee Assistance Program

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

1%

2 year growth

1%
Splitpoint Solutions
Jun 8th, 2026
AI Observability without unified context is just a new Blind Spot.

AI Observability without unified context is just a new Blind Spot. Why the rush of AI observability launches from Grafana, Datadog, and Dynatrace misses the deeper question, and how Mivu's unified-context approach answers it. What is AI observability, and why is a new dashboard not enough? AI observability is the practice of monitoring agent reasoning, tool calls, prompt-response quality, and decision health in production. A bolt-on agent dashboard cannot see why an agent fails when the underlying network, infrastructure, or application context shifts. Mivu's Unified Logic correlates agent telemetry with the full stack so root cause is one query away, not three tools away. Across the last 30 days, the three loudest names in observability (Grafana Labs, Datadog, and Dynatrace) have all moved aggressively into AI observability, LLM monitoring, and agent telemetry. The framing is the same. AI workloads are a blind spot. The fix being sold is also the same: a new dashboard, a new SDK, a new benchmark. The deeper truth is harder. Predictive alerts, data latency, and network reliability, the semantic siblings of real observability, still decide whether an AI agent answers correctly or melts down at 2 a.m. The week the industry agreed on the problem. In a single news cycle, the three competitors most often raised in Mivu sales conversations have converged on the same narrative. * Grafana Labs announced AI Observability in Grafana Cloud at GrafanaCON 2026 in Barcelona on 21 April 2026, framed publicly as closing the AI Blind Spot. Alongside it, the company open-sourced o11y-bench, a benchmark for AI agents running observability workflows, and shipped a new CLI, GCX, aimed at engineers working inside AI-assisted IDEs. * Datadog spent its Q1 2026 earnings cycle (results posted 7 May 2026) reinforcing its LLM Observability, Bits AI Security Analyst, GPU Monitoring, and Experiments line-up, and opened registration for DASH 2026 in New York on 9 and 10 June. AI workloads are now Datadog's headline growth narrative, not a side bet. * Dynatrace used its Perform 2026 cycle to push Davis AI further: root-cause analysis, automated remediation workflows, natural-language explanations, and new multi-cloud integrations across AWS, Azure, and GCP, all packaged as autonomous intelligence and preventive operations. Three different launch decks, and one shared message: AI agents in production need their own observability layer. It is the right diagnosis. It is the wrong cure. The pivot: the AI Blind Spot is really an AI Context Gap. A modern AI agent is, at minimum, a chain that includes a language model, a vector store, several tool integrations, a network path to each tool, the infrastructure each tool sits on, and the application the user actually interacts with. When the agent fails, the failure is almost never in the prompt-response cycle alone. It is usually upstream, in a slow database, a packet-loss spike on the egress path, a stale cache, or an overloaded GPU node. Treating agent telemetry as a separate silo, namely a dedicated AI observability product alongside your existing logs, metrics, and traces, recreates the very problem observability was supposed to solve. You replace one set of dashboards that engineers had to mentally correlate with a new, slightly more expensive set. The blind spot did not move. It got renamed. This is the Pivot Matrix in motion. * Where competitors say AI Observability, Mivu says Unified Logic. Agent traces, infrastructure telemetry, and network signals sit on one timeline because the failure crosses all three. * Where competitors say open benchmarks for agents, Mivu says proven outcomes in production. Benchmarks measure tools. Deployments measure value. * Where competitors say autonomous remediation, Mivu says Predictive Scale. Anomaly detection on the infrastructure context an agent depends on makes the remediation step rarer in the first place. What Splitpoint Solutions see in real deployments. In its deployment with enterprise clients running mixed AI and traditional workloads, the pattern is consistent. The most expensive AI-agent incidents are not model-quality failures. They are context failures that the model could not have known about. A retrieval-augmented agent goes silent because a vector-store sidecar lost its lease. A customer-support copilot starts hallucinating dates because an internal time-sync service drifted. A code-review agent times out because the GitOps webhook is being throttled by a misconfigured WAF rule. Its lead engineers recommend three operating habits for any team putting AI agents into production. * Instrument the path, not just the prompt. Capture network latency, infrastructure health, and application response codes on every tool call the agent makes. * Treat agent sessions as first-class telemetry, but in a unified stream. Grafana is right that conversations are telemetry. They are wrong that the conversations belong in their own warehouse. * Make "who broke whom" answerable in one query. If your engineers have to swivel between an AI dashboard, an APM tool, and a NetFlow viewer to triage an agent incident, your observability is fragmented, regardless of how many AI features are on the marketing page. | Mivu Laboratory Insight (engineering to verify exact figures before publish) In Mivu deployments correlating AI-agent telemetry with the underlying network and infrastructure context, an estimated majority of agent-level incidents have an upstream infrastructure or network trigger that pre-dates the agent error by tens of seconds to minutes. That is precisely the window a predictive, unified-context platform is built to surface. | Mivu compared to a generic AI-observability bolt-on. | Dimension | Generic AI-Observability Bolt-on | Mivu Unified Logic | | Signal model | Agent traces and prompts treated as a separate silo, alongside existing logs and metrics. Agent signals correlated with network, infrastructure, and application telemetry on one timeline. | | Failure mode it catches | Content-quality or prompt-level errors: the model said the wrong thing. Context-level root cause: the model said the wrong thing because upstream packet loss spiked. | | Posture | Reactive. Alerts fire after an agent run fails. Predictive. Anomaly detection covers the infrastructure context the agent depends on. | | Deployment footprint | Yet another collector, SDK, and dashboard to maintain. Lightweight probes fed into the same Mivu fabric you already run. | | Measure of success | Benchmark score on an open evaluation suite. Business outcomes in production: incidents avoided, MTTR, customer-impact minutes. | How Mivu operationalises Unified Logic. Mivu is built as an observability fabric. Lightweight probes feed network, infrastructure, and application telemetry into a single correlation engine, with predictive analytics layered on top. Adding AI-agent telemetry to that fabric is not a separate product. It is the next signal type on a timeline engineers already trust. Explore the Mivu monitoring stack, or see how the same fabric carries network observability for South African enterprises at splitpoint.io/sa/network-monitoring. Frequently asked questions. Both. The category is real because multi-agent systems, tool-using LLMs, and retrieval-augmented workflows generate telemetry that traditional APM tools were not designed to surface. The marketing rebranding is also real, because much of what is being sold as net-new is repackaged trace and log infrastructure. The question to ask a vendor is whether agent signals share a timeline with the rest of your stack or live in their own warehouse. How does Mivu compare to Grafana, Datadog, or Dynatrace for AI workloads? Mivu does not compete on the number of AI features in a marketing brochure. Splitpoint Solutions compete on time-to-root-cause when an AI workload misbehaves in production. Because Mivu unifies network, infrastructure, and application telemetry on one fabric, agent signals slot in as a new layer, not a new tool. For most enterprise teams that reduces tool sprawl rather than adding to it. What should an enterprise team do this quarter? Three steps. First, audit the AI agents you have in production today and list every external dependency each one touches. Second, pick one agent and shadow it with unified telemetry for a fortnight to see how often agent-level incidents are actually upstream incidents. Third, if the answer is often, consolidate observability before adding another AI-specific dashboard. Book a 30-minute Mivu Unified observability demo. See how lightweight probes, predictive analytics, and a single correlation layer change the way AI-agent incidents are triaged. Book a demo here.

Business Wire
Apr 8th, 2026
Grafana Labs expands AI-powered observability in Latin America with Santiago event

Grafana Labs is hosting an Observability Sessions event in Santiago on 15 April, bringing AI-powered observability solutions to Latin America. The company has grown its local team by 30% over two years to meet rising demand from regional organisations including LATAM Airlines, Casas Bahia and Hona. The event will feature technical sessions and demonstrations of Grafana Assistant, now generally available in Grafana Cloud, which allows users to query observability data in plain language. Assistant Investigations, currently in public preview, acts as an autonomous agent coordinating across metrics, logs, traces and profiles to identify root causes during incidents. According to Grafana Labs' 2026 Observability Survey, over 61% of South American respondents identified root cause analysis as AI's highest potential value area in observability.

IT Security News
Apr 7th, 2026
GrafanaGhost: attackers can abuse Grafana to leak Enterprise data.

GrafanaGhost: attackers can abuse Grafana to leak Enterprise data. 2026-04-07 16:04 By targeting Grafana's AI components, attackers can point to external resources and inject indirect prompts to bypass safeguards. Read the original article: Grafana Labs has disclosed a critical security vulnerability affecting Grafana Enterprise that could allow attackers to escalate privileges and impersonate users. The flaw, tracked as CVE-2025-41115, has received the maximum CVSS score of 10.0, making it one of the most severe vulnerabilities discovered in recent times. The vulnerability exists in Grafana's... November 21, 2025 In "Cyber Security News" Computer Security Grafana Labs has released critical security patches addressing a severe vulnerability in its SCIM provisioning feature that could allow attackers to escalate privileges or impersonate users. The flaw, tracked as CVE-2025-41115 with a CVSS score of 10.0 (Critical), affects Grafana Enterprise versions 12.0.0 through 12.2.1 under specific configurations. Organizations using... November 21, 2025 A high-severity cross-site scripting (XSS) vulnerability in Grafana could allow attackers to redirect users to malicious websites. The vulnerability, tracked as CVE-2025-4123 received a CVSS score of 7.6 (HIGH), allows attackers to exploit client path traversal and open redirect to execute arbitrary JavaScript code through custom frontend plugins. The vulnerability... May 22, 2025 In "Cyber Security News"

Nurbak
Apr 2nd, 2026
New Relic vs Grafana: which monitoring stack in 2026?

New Relic vs Grafana: which monitoring stack in 2026? An honest comparison of New Relic (managed SaaS, per-user pricing) and Grafana (open-source, self-hosted or cloud). Pricing, features, learning curve, and when neither fits your needs. New Relic and Grafana represent two fundamentally different approaches to monitoring. New Relic says: "Here is a complete platform. Send us your data and we handle everything." Grafana says: "Here are the building blocks. Assemble the stack that fits your needs." Both approaches work. The question is which one fits your team size, budget, technical capacity, and tolerance for operational overhead. New Relic: the Managed SaaS platform. New Relic is a fully managed observability platform. You install their agent, it collects metrics, traces, logs, and errors, and everything appears in a single web interface. No infrastructure to manage. No databases to run. No configuration files to maintain. What you get. * APM (Application Performance Monitoring): Auto-instrumentation for most languages and frameworks. Transaction traces, slow query analysis, error analytics. * Infrastructure monitoring: Host metrics, container metrics, Kubernetes monitoring. Integrations with 750+ technologies. * Log management: Ingest, search, and analyze logs. Correlate logs with traces and errors. * Distributed tracing: End-to-end request tracing across services. * Synthetics: Uptime monitoring with scripted browser checks. * Alerts: Threshold-based and anomaly detection alerting with PagerDuty, Slack, email integrations. * NRQL (New Relic Query Language): SQL-like language for querying all your telemetry data. Powerful but proprietary. Pricing (2026). New Relic changed to a user-based pricing model: * Free tier: 1 full-platform user, 100GB/month data ingest, forever free. * Standard: $49/user/month, 100GB free then $0.35/GB. * Pro: $349/user/month, advanced features. * Enterprise: Custom pricing, HIPAA compliance, SSO. The per-user model is a double-edged sword. For a solo developer or small team of 2-3, the free tier is genuinely generous - 100GB is enough for most applications. For a team of 15 engineers, the cost is $735/month on Standard before data charges. That adds up fast. Strengths. * Zero infrastructure to manage. Install agent, see data. * One platform for everything. No tool integration headaches. * NRQL is genuinely powerful for ad-hoc queries. * Free tier is production-ready, not a trial. Weaknesses. * Per-user pricing gets expensive for growing teams. * Vendor lock-in. NRQL, custom instrumentation, dashboards - all proprietary. * Data ingest costs are unpredictable. A noisy microservice can blow your budget. * UI can be overwhelming. There are so many features that finding what you need takes time. Grafana: the open-source stack. Grafana is not a single product - it is a ecosystem. At its core, Grafana is a visualization and dashboarding tool. But a complete monitoring stack typically includes: * Grafana: Dashboards, alerting, and visualization. * Prometheus: Metrics collection and storage (time-series database). * Loki: Log aggregation (like a lightweight ELK). * Tempo: Distributed tracing (stores traces in object storage). * Mimir: Long-term metrics storage (Prometheus-compatible). * Alloy (formerly Grafana Agent): Telemetry collector that ships data to all of the above. Self-hosted vs Grafana Cloud. Self-hosted: All components are open-source. You run them on your own infrastructure. Free, but you are responsible for uptime, scaling, backups, and upgrades. Grafana Cloud: Managed version of the entire stack. Free tier includes: * 10,000 active metrics series * 50GB logs/month * 50GB traces/month * 500 VUH (virtual user hours) for k6 load testing * 50GB profiles/month Paid plans start at $29/month for Grafana Cloud Pro, scaling based on usage. * No vendor lock-in. Prometheus, OpenTelemetry, and PromQL are industry standards. * Extremely flexible. You can build exactly the stack you need. * Beautiful dashboards. Grafana's visualization is best-in-class. * Massive community. Thousands of pre-built dashboards, exporters, and integrations. * Cost-efficient at scale. Open-source components mean you pay for infrastructure, not licenses. * Operational overhead. Running Prometheus + Loki + Tempo + Grafana is a lot of infrastructure. * Steeper learning curve. PromQL, LogQL, and TraceQL are three different query languages. * Assembly required. New Relic works out of the box. Grafana requires configuration, integration, and ongoing maintenance. * Alerting is functional but not as sophisticated as dedicated tools like PagerDuty or OpsGenie. Head-to-Head comparison. | Dimension | New Relic | Grafana (Stack) | | Type | Managed SaaS | Open-source / Managed Cloud | | Setup time | Minutes (install agent) | Hours to days (self-hosted) / Minutes (Cloud) | | Infrastructure management | None | Significant (self-hosted) / None (Cloud) | | APM | Built-in, auto-instrumented | Via Tempo + OpenTelemetry | | Logs | Built-in | Loki | | Dashboards | Good | Best-in-class | | Query language | NRQL (proprietary) | PromQL, LogQL, TraceQL (open standards) | | Vendor lock-in | High | Low (open standards) | | Free tier | 1 user, 100GB/month | 10K metrics series, 50GB logs/month | | Paid pricing | Per user ($49+/user/month) | Per usage (metrics, logs, traces) | | Best for | Teams wanting zero ops | Teams wanting flexibility and control | When to choose New Relic. * Your team has 1-3 engineers and you do not want to manage monitoring infrastructure. * You need APM, logs, and infrastructure in one place with zero setup. * You value simplicity over flexibility. * Your data volume is under 100GB/month (free tier is genuinely useful). When to choose Grafana. * You already run Kubernetes and your team is comfortable with Prometheus. * You want to avoid vendor lock-in and use open standards. * You need highly customized dashboards and visualizations. * You have the DevOps capacity to manage the stack (or use Grafana Cloud). * You are cost-sensitive at scale - open-source scales cheaper than per-user pricing. When neither fits. Both New Relic and Grafana are designed for teams operating infrastructure. They assume you have servers, containers, or at least a multi-service architecture worth monitoring. But many Next.js applications are deployed on Vercel or similar platforms where you do not manage infrastructure. You do not have hosts to monitor. You do not have Prometheus endpoints to scrape. What you have is API routes that need to be fast, reliable, and monitored. For this scenario: * New Relic's agent-based approach does not work on serverless without significant configuration. * Grafana's Prometheus-based stack has nothing to scrape in a serverless environment. Nurbak Watch is built for this gap. It runs inside your Next.js server via instrumentation.ts - five lines of code - and monitors every API route from the inside. No agents, no exporters, no infrastructure. Alerts via Slack, email, or WhatsApp in under 10 seconds. $29/month flat, free during beta. If you grow into managing your own infrastructure, New Relic or Grafana will be there. Start with what your architecture actually needs. The Nurbak Team builds developer-first API monitoring tools. Nurbak share insights on uptime, performance, alerting, and best practices for keeping APIs healthy in production. Ready to try it? Nurbak Watch is free during beta. 5 lines of code. First alert in under 5 minutes. Comparisons

GameFabrique
Mar 26th, 2026
Eliminating static waste: automating capacity management with Dynamic Buffers.

Eliminating static waste: automating capacity management with Dynamic Buffers. * March 26, 2026 * Samuel Good As your multiplayer game scales, relying solely on fixed buffer sizes to manage unpredictable player spikes can turn off-peak hours into a costly operational blind spot. Reserving permanent spare cloud capacity for these sudden surges means paying for idle compute when player traffic drops. This provisioning model inflates your infrastructure budget, effectively acting as a tax for phantom workloads that sit empty waiting for traffic. GameFabric eliminates this idle-capacity tax through intelligent, automated orchestration. GameFabric operates on a synergistic hybrid model: your predictable player concurrency runs on highly performant, cost-efficient bare metal, while unpredictable player surges burst automatically into the cloud. Dynamic Buffers enforce this philosophy by intelligently scaling your elastic cloud servers while keeping your core bare metal servers ready. By concentrating scaling actions strictly on the cloud tier, GameFabric preserve the dedicated nature of your bare metal foundation. To achieve true cost optimization without compromising the player experience, infrastructure requires intelligent, demand-based scaling. Dynamic Buffers replace static waste with automated capacity management, dynamically growing and shrinking your buffer sizes in direct correlation with player demand. Dynamic Buffering is not a blunt, centralized trigger; it's a sophisticated approach designed specifically for autonomous container orchestration within your Armada and ArmadaSet deployments. * Cluster-Level Autonomy: Unlike autoscalers that rely on a central API, GameFabric's dynamic scaling logic operates directly inside your game cluster. This localized architecture removes single points of failure. Even in the event of an API disruption, your game servers continue scaling flawlessly to meet real-time player demand. * The Fallback Safety Net: Dynamic Buffers are underpinned by a secure 'floor.' Whether you define a static manual buffer or allow its system to calculate a safe baseline, this safety net ensures your game remains highly available even if the dynamic system is overridden. While the system automates capacity management, your engineering team retains absolute control over scaling behavior. Through the GameFabric UI, LiveOps teams can use a simple slider to dictate the system's priorities: * Cost Efficient: Configures a leaner infrastructure to maximize savings during stable periods. * Availability: Configures the orchestrator to scale up fast and scale down slow, maintaining a robust buffer to handle massive player influxes safely and ensuring players aren't left waiting for new servers to spin up. Effective fleet management demands clear visibility into system behavior. Because GameFabric integrates natively with Prometheus and Grafana, your LiveOps team can utilize customized dashboards to visualize buffer adjustments over time, correlating scaling events directly with Concurrent User (CCU) demand. Should you need to lock your infrastructure for a specific live event, the system features a manual override. Disabling Dynamic Buffers instantly halts automated tuning, safely reverting the fleet to your statically defined fallback buffer without interrupting active game sessions. Reclaim your cloud budget and offload the operational overhead of manual fleet management to an orchestrator built specifically for the realities of live-service gaming. Reach out today for your personalized demo to see automated capacity management in action.

INACTIVE