Gremlin, whose tools help app and online service providers simulate failure scenarios like storage error, database congestion, and outages, raises $18M Series B
“Slack is down.” It's a headline we have had blaring at TechCrunch on numerous occasions (mostly because we actually get work done …
Context & Ripple Effects
Gremlin's $18M Series B lands in a year when outage pain was already forcing buyers to act: Slack had logged 40 days with outages since April 2017 and responded by building a dedicated safety engineering team to reduce them. Gremlin's pitch is the other half of that equation — instead of waiting for failures, its tools inject storage errors, database congestion, and outages on purpose so providers find their breaking points first.
The raise also sits early in what became a funding arc across the reliability stack: three years later, Cribl pulled a $200M Series C for infrastructure data monitoring at a reported $1.5B valuation, and Gluware raised $43M led by Bain Capital for network-outage prevention orchestration.
First-order effects
- Gremlin gets growth capital to scale failure-simulation tooling for app and online service providers, giving engineering teams a way to rehearse outages rather than discover them in production.
Second-order effects
- Vendors on adjacent layers of the same stack are pushed to define their boundary with chaos testing — Cribl's monitoring and Gluware's prevention-orchestration tools either complement a simulate-first workflow or compete for the same reliability budget line.
Third-order effects
- If the pattern holds, resilience shifts from an internal ops discipline to a procured multi-vendor toolchain — simulate, monitor, prevent — with each layer attracting its own venture-backed specialist, as the Cribl and Gluware rounds suggest.
The trend: Infrastructure reliability is becoming a funded, layered software market where companies buy failure simulation, monitoring, and prevention as separate products rather than building resilience in-house.