Kolton Andrus once helped keep Amazon’s retail site online. Later he built fault-injection tools at Netflix after the company released Chaos Monkey. Now his company sells the idea that systems improve when deliberately broken. On October 7, Gremlin made that process run itself.
The firm launched Foresight AI. Agents installed in customer environments inject latency, kill pods, exhaust memory. A cloud service watches telemetry. When something cracks, the system identifies root cause, suggests a code or configuration fix, applies or recommends it, then repeats the test until the weakness disappears. Humans stay in the loop for approval at key moments. But the loop runs without constant engineer attention.
Gremlin calls its secret sauce the Failure Atlas. This repository holds results from millions of experiments conducted over the past decade across tens of thousands of systems. The language model does not guess from internet scraps. It reasons from observed cause and effect. Hallucinations drop. Answers stay grounded. Gremlin’s own announcement describes the Atlas as the only record of its kind.
Andrus pulls no punches. “A lot of AI solutions are, ‘Hey, we took a guess. Here you go. Good luck,’” he told LavX News. “We actually induce the failure, and then we tell you how to prevent it.”
Most tools react. They wait for an incident, then attempt explanation. Foresight acts first. It scans for potential failure modes before alerts fire. It recommends fixes as infrastructure-as-code changes or configuration patches. It reruns the exact experiment that exposed the risk to confirm resolution. Reliability scores track progress across services and teams. Leaders gain numbers to guide investment.
The timing feels deliberate. AI systems generate code at high velocity. They introduce unknown dependencies, misconfigured scaling, broken redundancies. Development outruns traditional testing. Observability captures problems after they bite production. Andrus argues this gap widens. “AI-driven development means shipping code at 10X velocity; it also means 10x the opportunity for bugs, risks, and failures,” he said in a Help Net Security release.
Gremlin built on its original fault-injection platform. Agents already lived in customer infrastructure. The new layer adds analysis, experiment generation, root-cause reasoning, remediation, and verification. A technical project manager agent tracks commitments, produces weekly reports, sends Slack updates. The product mixes closed and open models, swapping whichever performs best at any moment.
Early customers see the difference. Teams that once lacked time or expertise for consistent chaos experiments now receive automated recommendations. Fixes get validated against the original test. Scores make reliability visible and comparable. One investor noted the shift. “Gremlin has always tried to help engineering teams improve the resiliency of their applications by running experiments proactively to identify weaknesses in their distributed systems, but many teams didn’t have the time or expertise to do that consistently,” said Mike Dauber, general partner at Amplify Partners, quoted in Gremlin’s PR Newswire announcement.
Yet questions remain. Autonomous fault injection carries risk even with blast-radius controls and halt conditions. Gremlin runs in regulated environments and holds SOC 2 Type II certification. It offers a Private Edition for isolated networks and uses AWS PrivateLink so data stays inside the customer’s cloud. Still, handing agents permission to break production-like systems demands trust.
The company’s history helps. Founded by engineers from Amazon and Netflix, Gremlin commercialized chaos engineering for enterprises. It moved beyond one-off experiments into continuous reliability management. Foresight represents the latest step. The agentic loop analyzes, tests, explains, remediates, verifies. It repeats. Each cycle draws on the Failure Atlas for precision.
Industry observers point to a broader pattern. AI accelerates change. Traditional SRE practices struggle to keep pace. Reactive tools clean up afterward. Proactive systems that find weaknesses first gain favor. TFiR’s interview with Andrus frames the issue clearly: AI ships more defects. Foresight finds, fixes, and verifies before incidents occur.
Gremlin positions Foresight as an add-on to existing services. Pricing details remain limited in initial coverage, but the company emphasizes measurable outcomes. Reliability scores give every team a shared standard. Dashboards show risk and progress over time. The goal stays practical. Let AI move fast. Keep systems standing.
Not every organization will hand over the keys immediately. Many will start with recommendations rather than automatic application. The product keeps humans at critical junctures. That design choice reflects lessons from a decade of real failures. The Atlas records not just what broke but what restored service. Recommendations flow from proven patterns rather than theory.
Competitors in observability and AIOps now face a different challenge. Their tools shine after the fact. Gremlin claims to act before. The distinction matters as distributed systems grow more complex and code changes arrive faster than ever. Whether Foresight delivers consistent results across varied environments will determine adoption. Early beta success suggests momentum.
Andrus and his team built the Failure Atlas through years of customer experiments. That data now powers agents that operate at AI speed. The loop closes faster than manual processes ever could. Failures get induced, analyzed, fixed, and verified in repeated cycles. Systems emerge stronger. Or so the pitch goes.
Engineers who remember manual chaos exercises recognize the appeal. Designing the right experiment, running it safely, interpreting results, implementing fixes, and retesting consumed hours. Foresight compresses much of that work. It does not eliminate human judgment. It multiplies it.
The Register captured the spirit in its headline. AI can fail faster than the average SRE. Gremlin simply decided to make that speed useful. The company turned its breaking business into an automated reliability engine. The rest of the industry will watch closely to see how far the approach travels.
Gremlin’s Foresight AI Turns Chaos Engineering Into an Automated Reliability Machine first appeared on Web and IT News.
