Richard Teachout // Teachout.com
← All writing

The Kill Switch You Hope You Never Need (and Must Still Build)

Richard Teachout
Richard Teachout CTO at Ashley Furniture Industries - Executive Tech Leader, Entrepreneur, AI leader, Architect, Problem Solver, Ex-Developer. September 29, 2026
AI
The Kill Switch You Hope You Never Need (and Must Still Build)

The kill switch is the most boring thing you'll build, and it's the one thing you can't ship without.

I've seen what happens when AI goes wrong in production. Not the dramatic version — the slow, expensive version. The model starts behaving differently, the outputs start getting weird, and the team's only option is to scramble. Pull the integration. Turn off the endpoint. Call the vendor. Every option takes time, and every minute the bad behavior runs, it's compounding — bad answers reaching customers, bad data polluting the pipeline, bad decisions being logged as if they were good.

The kill switch is the answer to that scramble. It's the pre-built, tested, boring path to stop the system before the damage compounds. You hope you never pull it. You're irresponsible if you ship without it.

What a kill switch actually is

The kill switch isn't one thing. It's three, and they cover different situations.

The soft stop: stop new requests from reaching the model. Incoming traffic pauses, the model stops working, and everything already in flight finishes. This is the everyday switch — the one you pull when the escalation rate spikes, when the drift alarm fires, when the review queue is drowning. Nothing is lost. The system just pauses while you look.

The rollback: revert to the last known-good version. The model changed, the prompt changed, the config changed — and the change is the problem. Rollback puts the previous version back. This is the switch for the most common failure mode: the upgrade that was worse than what it replaced.

The hard stop: kill everything, including in-flight work, and isolate the system. This is the emergency switch — the one for the case where continuing anything, even finishing what started, is worse than losing it. The hard stop is about blast radius. When you pull it, you're saying the cost of more output exceeds the cost of interrupted work.

The kill switch is three paths: pause, roll back, or stop hard. Each one is built, tested, and boring — so the moment you need it, it's the easiest thing you'll do all week.

The blast-radius design

The kill switch only works if it's scoped. The design question isn't just "can we stop the AI" — it's "can we stop the part that's bad without stopping everything else?"

The system has blast-radius boundaries. The customer-facing path is one radius. The internal-analysis path is another. The batch pipeline is a third. A good kill switch design has switches at each boundary, so the failure in the customer-facing path doesn't take down the internal tooling that's working fine. The team that can only stop everything is a team that won't stop anything, because stopping everything is too expensive.

The scoping comes from the architecture. Each integration point gets its own switch — the model call, the prompt template, the routing rule, the review queue. A switch at each point means you can stop the bad model without stopping the good one, stop the bad prompt without stopping the model. The switches are cheap to add at build time and impossible to add cleanly after the incident.

The testing that makes it real

A kill switch that's never been tested is a switch that doesn't exist. The first time you need it is not the time to find out it doesn't work, or that it takes an hour, or that the person who knows how to run it is on vacation.

The testing discipline is the same as for any emergency system: scheduled, boring, and treated as routine. Monthly, pull the soft stop and verify traffic pauses cleanly. Quarterly, exercise the rollback on a staging environment. Annually, do a full hard-stop drill, end to end, with the team that would actually run it.

The drills are also training. The person who has pulled the soft stop in a drill knows exactly what it feels like when the real thing happens. The team that has practiced the hard stop doesn't freeze at the moment of decision. The drill is what turns the kill switch from a feature into a capability.

The decision framework

The hardest part of the kill switch isn't building it. It's deciding to pull it. The decision has a framework, and it should be written down before the incident, not during it.

Pull the soft stop when the signals say the system is degrading — escalation rising, drift climbing, confidence dropping — even before the errors are visible to customers. The soft stop is cheap. Use it early.

Pull the rollback when a change is the likely culprit. New model, new prompt, new config, and the behavior changed — roll back first, investigate second. The rollback is the fastest way to restore service, and the investigation is easier on a working system.

Pull the hard stop when the damage is compounding faster than the diagnosis. The rare case where continuing is worse than stopping, and stopping everything is the only move that limits the blast radius. The framework is written down so the decision is made by the rule, not by the moment.

The kill switch is the most boring part of the AI system, and that's exactly why it's the most important. Build the three paths, scope them to the blast radius, test them on a schedule, and write down the decision framework before you need it. You hope you never pull it. But the day you need it, it has to be the easiest decision of your week.

Think this argument fits your event? Tell me about the room — the calendar is selective.

Start a conversation