Full-Time

Data Scientist

AI/ML

Gremlin

Gremlin

51-200 employees

Chaos engineering platform for reliability testing

Compensation Overview

$220k - $290k/yr

+ 401k Matching + Equity

No H1B Sponsorship

Remote in USA

Remote

Remote within the United States.

Category
Data & Analytics (1)
Required Skills
Agile
Data Science
Machine Learning
DevOps
Reinforcement Learning

Get referred to Gremlin

See people who can refer or advise you

Requirements
  • 5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development.
  • Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning
  • Experience building data pipelines and feature stores that support both offline training and real-time inference
  • Experience with agile development environments and practices
  • Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices
  • Comfort partnering with platform engineers and SREs to turn research into shipped product features
  • Strong at breaking down ambiguous problems into concrete actions and milestones
Responsibilities
  • Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems
  • Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments
  • Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior
  • Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference
  • Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product
  • Apply advanced techniques, including causal inference, graph ML, time-series modeling, and reinforcement learning, to continuously improve the accuracy and actionability of automated failure analysis
  • Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery
  • Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies
Desired Qualifications
  • Experience with chaos engineering, site reliability engineering, or distributed systems
  • Background in agentic AI systems or large-scale causal inference in production
  • Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure
  • Working in Remote first environments
  • Has been on-call and participated in an incident management program

Gremlin provides a platform that helps businesses, especially financial services firms, keep their software reliable and available by running controlled experiments that simulate failures (chaos engineering). The product works by letting teams design and run disruptions in their systems, measure how the software and infrastructure respond, and track changes over time, all through a subscription-based service with different tiers. What sets Gremlin apart is its industry focus on financial services, its emphasis on continuous testing of both applications and infrastructure, and its Enterprise Chaos Engineering Certification (GECEC) program, which trains and certifies professionals in chaos engineering and builds a community around reliability practices. The company’s goal is to reduce the risk of customer-facing outages and performance problems, helping customers meet high availability standards, protect revenue, and improve overall user experience.

Company Size

51-200

Company Stage

Series B

Total Funding

$26.8M

Headquarters

San Jose, California

Founded

2016

Get referred to Gremlin

See people who can refer or advise you

Simplify Jobs

Simplify's Take

What believers are saying

  • July 28, 2026 Dynatrace integration expands Gremlin inside an enterprise observability standard.
  • February 2026 disaster recovery testing targets cloud outage budgets after costly 2025 failures.
  • April 21, 2026 Carahsoft partnership unlocks government procurement across SEWP V and NASPO.

What critics are saying

  • Gremlin reported only 42 employees in April 2026, limiting enterprise support capacity.
  • No public 2026 funding signals force growth from subscriptions and partner channels.
  • Dynatrace, AWS, and observability vendors can bundle similar resilience testing, commoditizing Gremlin.

What makes Gremlin unique

  • Gremlin pairs chaos engineering with reliability scoring across observability workflows.
  • July 2026 Dynatrace app embeds testing, halt conditions, and reliability scores.
  • Carahsoft contract access opens Gremlin to federal, state, and local buyers.

Help us improve and share your feedback! Did you find this helpful?

Benefits

Health Insurance

Dental Insurance

Vision Insurance

401(k) Retirement Plan

401(k) Company Match

Company Equity

Growth & Insights and Company News

Headcount

6 month growth

0%

1 year growth

1%

2 year growth

0%
PR Newswire
Jul 28th, 2026
Gremlin launches native Dynatrace app for resilience testing and reliability scoring

Gremlin has launched a native app for Dynatrace that enables engineering teams to run resilience tests and view reliability scores directly within the observability platform. The app combines observability with reliability testing, allowing teams to test services using existing Dynatrace alerts and metrics as health checks. Teams can run full test suites or specific tests with automated halt conditions, view results cross-referenced against key Dynatrace metrics like response time and failure rate, and integrate Gremlin reliability scores into Dynatrace dashboards. The app supports continuous reliability testing, disaster recovery validation, and risk detection across bare metal, on-premise, multi-cloud, and serverless environments. Gremlin is trusted by global enterprises including four of the five largest US banks. Customers have achieved a 50% reduction in downtime and 90% reduction in disaster recovery testing time. The app is available today in the Dynatrace Hub.

USA News Hour
Jul 28th, 2026
Gremlin launches native app for Dynatrace, adding resilience testing and reliability scoring to the observability platform.

Gremlin launches native app for Dynatrace, adding resilience testing and reliability scoring to the observability platform. New app lets engineering teams run Gremlin resilience tests and surface forward-looking reliability scores directly within Dynatrace using the metrics and alerts they already trust. SAN FRANCISCO, July 28, 2026 /PRNewswire/ - Gremlin, the reliability management platform trusted by the world's largest enterprises, today announced the Gremlin app for Dynatrace. The app lets engineering teams run Gremlin resilience tests, see their impact on systems, and track forward-looking reliability scores directly within Dynatrace, extending the world-class observability platform's capabilities so teams can also validate and measure the resilience of every service. An effective reliability practice is built on two complementary disciplines: observability, which gives teams real-time visibility into system performance, and reliability testing, which validates how those systems will respond to failures. The Gremlin app for Dynatrace brings these together into a single workspace, using the Dynatrace metrics and alerts teams already trust as the foundation for proactive reliability testing. With the app, teams can: * Test services within Dynatrace: run a service's full test suite or specific tests using existing Dynatrace alerts and events as health checks and automated halt conditions to stop tests if services move outside defined thresholds. * See results using trusted metrics: every test result is cross-referenced against key Dynatrace metrics, including response time, requests per minute, and failure rate, with a full pass/fail history for each service. * Bring reliability scores into existing dashboards: Gremlin reliability scores, test results, and pass rates appear in Dynatrace dashboards, so forward-looking reliability data joins the view executives and stakeholders already trust. The app builds on Gremlin's existing Dynatrace integration, adding a native user interface and reliability scores within Dynatrace dashboards. It supports continuous reliability testing, disaster recovery validation, risk detection, dependency mapping, and reliability reporting and governance. Gremlin's production safety controls - blast radius management, halt conditions, and automatic rollback - apply throughout, and the app works across bare metal, on-prem, multi-cloud, and serverless environments. "Dynatrace shows teams exactly how their systems are performing," said Kolton Andrus, CEO and Co-founder of Gremlin. "Gremlin takes that same trusted data and uses it to prove how those systems will perform under failure, then turns it into a reliability score teams can track over time.Together, they give teams the full picture: what's happening now, and what will happen when something breaks." "Our customers rely on Dynatrace as their source of truth for the health of their systems," said Philippe Deblois, Global Vice President, Solutions Engineering at Dynatrace. "Bringing Gremlin's best-in-class enterprise reliability testing and scoring into that environment is a natural extension of the value our platform delivers. It gives teams a way to validate resilience and act on forward-looking insight, all powered by the Dynatrace data and alerts they already depend on every day." Gremlin is trusted by global enterprises across financial services, SaaS, retail, and media, including 4 of the 5 largest US banks. Customers using Gremlin have achieved outcomes including a 50% reduction in downtime, a 90% reduction in disaster recovery testing time, and 99.99% availability on critical platforms. The Gremlin app for Dynatrace is available today in the Dynatrace Hub. To learn more or request a demo, visit gremlin.com/demo. About Gremlin Gremlin is the reliability management platform trusted by the world's largest enterprises across financial services, SaaS, retail, and media. Combining failure testing, passive risk detection, and dependency mapping, Gremlin gives engineering teams the predictive data they need to systematically measure, manage, and improve their reliability. Learn more at gremlin.com. Media Contact SOURCE Gremlin Inc. Disclaimer: The above press release comes to you under an arrangement with PR Newswire. USA Newshour takes no editorial responsibility for the same. PR Newswire is a distributor of press releases headquartered in New York City.

PR Newswire
Jun 2nd, 2026
Gremlin launches no-code Failure Flags to test app reliability without changing source code

Gremlin, a leader in enterprise reliability management, has launched Failure Flags, a no-code solution enabling teams to test and improve application reliability without modifying source code. The technology works by proxying application network traffic through a dedicated container, eliminating the need for SDKs or code changes. Failure Flags supports serverless, container and hybrid environments including AWS Lambda, Azure Functions, Google Cloud Functions and Kubernetes. The system conducts targeted reliability experiments such as latency spikes and dropped packets, with built-in health checks that automatically halt tests if anomalies appear. The platform's reliability scoring feature tracks risk across both infrastructure and application layers. Founder Kolton Andrus said the approach allows engineers to quickly deploy and run experiments throughout the software development lifecycle, from cloud region outages to specific function call failures.

Carahsoft
Apr 21st, 2026
Gremlin and Carahsoft partner to bring reliability testing and Chaos Engineering to the Public Sector.

Gremlin and Carahsoft partner to bring reliability testing and Chaos Engineering to the Public Sector. Partnership makes forward-looking reliability management and disaster recovery testing technology easily accessible to government agencies. SAN JOSE, Calif., and RESTON, Va. - April 21, 2026 - Gremlin, the enterprise resilience testing and reliability management platform, and Carahsoft Technology Corp., The Trusted Government IT Solutions Provider(R), today announced a partnership. Under the agreement, Carahsoft will serve as Gremlin's Master Government Aggregator(R), making the company's reliability platform available to the Public Sector through Carahsoft's reseller partners and NASA Solutions for Enterprise-Wide Procurement (SEWP) V, Information Technology Enterprise Solutions - Software 2 (ITES-SW2), National Association of State Procurement Officials (NASPO) ValuePoint and OMNIA Partners contracts. "Partnering with Carahsoft is an important next step for Gremlin as we expand our reach into the Public Sector," said Dave Coughlin, Vice President of Sales at Gremlin. "Government agencies invest heavily in resilience and disaster recovery, but most have no systematic way to know if any of it actually works until something breaks in production. Carahsoft's trusted relationships across Government agencies remove traditional procurement barriers and make it easier for IT leaders to change that. Together, we're giving Federal, State and Local agencies a faster path to mission-critical resilience, helping them modernize with confidence, reduce risk and ensure the systems U.S. citizens rely on are always available." Public Sector organizations depend on digital systems to deliver essential services. These systems are increasingly complex, distributed and interconnected, which makes hidden reliability risks more likely. Misconfigured autoscaling, unknown dependencies, untested resilience mechanisms and non-compliant architecture can lead to outages that disrupt services and erode public trust. Gremlin's reliability platform uses Chaos Engineering principles to proactively uncover these risks before they cause real-world failures. By safely introducing controlled stress and failure scenarios, engineering teams can identify weaknesses, validate disaster recovery processes and strengthen system resilience without impacting production users. Key capabilities of the Gremlin platform include: * Forward-looking reliability measurement: Service-level reliability scores based on active failure testing and passive risk detection replace lagging indicators like uptime with a forward-looking view of where systems are at risk right now. * Combined active and passive assessment: Fault injection validates whether resilience mechanisms actually work, while passive detection surfaces configuration drift and deviate from standards, giving a more complete reliability picture than either approach alone. * Organization-wide reliability standardization and benchmarking: Standardized test suites and reliability scoring let leaders define what "good" looks like, measure every service against that standard, compare across teams and report progress in terms the organization understands. * Expertise-driven remediation guidance: Specific, actionable recommendations built on Gremlin's deep experience with the world's largest enterprises tell teams exactly what to fix, closing the gap between identifying risk and resolving it. * Measurable, provable resilience improvement: Continuous score tracking and executive-ready reporting let agencies demonstrate that reliability investments are actually reducing risk, turning a faith-based program into a data-driven one. "Carahsoft's new partnership with Gremlin brings advanced reliability capabilities to the Public Sector," said Natalie Gregory, Vice President for Open Source and DevSecOps Solutions at Carahsoft. "As agencies continue to modernize and rely on increasingly complex digital systems, resilience and availability are more critical than ever. By adding Gremlin's reliability management platform to our solution portfolio, Carahsoft and our reseller partners are enabling Federal, State and Local customers to identify risk earlier, strengthen mission-critical systems and deliver more reliable services to the constituents they serve." Gremlin's reliability platform is available through Carahsoft's SEWP V contracts NNG15SC03B and NNG15SC27B, ITES-SW2 Contract W52P1J-20-D-0042, NASPO ValuePoint Master Agreement #AR2472 and OMNIA Partners Contract #R240303. For more information, contact the Carahsoft Team at (703) 581-6680 or [email protected]; or read this Gremlin for Government Capability Statement to learn more about how Gremlin helps keep Government systems performant and highly available. About Gremlin Gremlin is the reliability management platform trusted by the world's largest enterprises across financial services, SaaS, retail, media and government. Combining failure testing, passive risk detection and dependency mapping, Gremlin gives engineering teams the forward-looking data they need to systematically measure, manage and improve reliability. Learn more at gremlin.com. About Carahsoft's AI Portfolio Carahsoft's Artificial Intelligence (AI) Portfolio includes leading and emerging technology vendors who are enabling Government agencies and systems integrators to harness the power of AI and ultimately meet mission needs; from creating efficiencies within agencies to bolstering national security and defense. Supported by dedicated AI product specialists and an extensive ecosystem of resellers, integrators and service providers, Carahsoft Technology Corp. help organizations identify the right technology for unique environments and provide access to technology solutions through its broad portfolio of contract vehicles. Its AI portfolio spans solutions for AI Infrastructure, Generative and Agentic AI, Autonomous Systems & Robotics and more. Learn more about Carahsoft's AI Solutions for Government here. About Carahsoft Carahsoft Technology Corp. is The Trusted Government IT Solutions Provider, supporting Public Sector organizations across Federal, State and Local Government agencies and Education and Healthcare markets. As the Master Government Aggregator for its vendor partners, Carahsoft Technology Corp. deliver solutions for Artificial Intelligence, Cybersecurity, MultiCloud, DevSecOps, Customer Experience and Engagement, Open Source and more. Working with resellers, systems integrators and consultants, its sales and marketing teams provide industry leading IT products, services and training through hundreds of contract vehicles. Visit Carahsoft Technology Corp. at www.carahsoft.com.

PR Newswire
Feb 3rd, 2026
Gremlin launches disaster recovery testing to help businesses prepare for cloud outages

Gremlin, a proactive reliability platform, has launched Disaster Recovery Testing to help businesses test zone, region and datacenter evacuations and failovers. The product enables companies to simulate complex disaster scenarios and validate failover systems with minimal engineering effort. The launch follows multiple high-profile cloud outages in 2025, including one affecting 70,000 companies with estimated losses of $581 million. Gremlin's platform allows organisations to conduct company-wide testing from a central command centre, with automated health checks and detailed reliability reports. The company works with dozens of Fortune 1000 clients, including four of the top five US banks. Gremlin's reporting capabilities also support companies preparing S-1 filings for the SEC and help public companies create annual 10-K reports detailing business operations and risks.