Richard Teachout // Teachout.com
← All writing

Your First AI Incident: A Post-Mortem Template

Richard Teachout
Richard Teachout CTO at Ashley Furniture Industries - Executive Tech Leader, Entrepreneur, AI leader, Architect, Problem Solver, Ex-Developer. October 5, 2026
AI
Your First AI Incident: A Post-Mortem Template

Your AI will have an incident. It's not a question of if. The only question is whether the incident teaches you something or just costs you something.

The AI incident looks different from the classic outage. The system wasn't down. It was up the whole time, responding, generating — wrong. Maybe for hours, maybe for days, before anyone noticed. The database outage is loud. The AI incident is quiet, and quiet incidents are the expensive ones, because they run longer before anyone sees them.

The post-mortem is how you turn the quiet incident into the lesson. The template below is the one I use, and it's built for the AI incident's specific shape: not "what broke" but "what was wrong, how long, and why didn't we see it sooner."

The five sections

Section one: what it was supposed to do. One paragraph, no drama. The task, the expected behavior, the quality bar it was built to meet. This section exists so the post-mortem starts from agreement — what the system was for — instead of diving straight into blame.

Section two: what it actually did. The behavior, the duration, the scope. How many outputs, over what period, affecting what. This is where the quiet nature of AI incidents shows up — the section often starts with "we don't know exactly when it started" and the fix for that gap becomes part of the action items.

Section three: how we found out. The detection story. Was it a customer complaint? A reviewer who noticed something off? A metric that finally crossed a line? This section is the most useful one, because it tells you how your detection failed. The incident that ran for three days wasn't a model failure. It was a detection failure, and this section names it.

Section four: why it happened. The root cause, written without blame. The model changed. The data drifted. The prompt drifted. The threshold was set wrong. The eval suite missed the case. Name the cause, and name the contributing factors — the missing signal, the unmonitored metric, the review queue that was too slow to notice.

Section five: what ships next. The action items, each one concrete, owned, and dated. Not "improve monitoring." "Add the override-rate signal to the dashboard by Friday, owned by the operator." The post-mortem's only job is to produce this section, and the section is only real if the items are specific.

The post-mortem's job isn't to explain the incident. It's to make the next one cheaper — shorter, smaller, and caught sooner. Five sections, no blame, and action items that ship.

The review cadence that makes it stick

A post-mortem that gets written and filed is a document. A post-mortem that gets reviewed is a system. The difference is the cadence.

The immediate review happens while the incident is fresh — the team walks the five sections, argues about the why, and assigns the items. Then the follow-up, thirty days later: did the items ship? Did the detection improve? Is the same incident possible tomorrow? The thirty-day check is where post-mortems usually die — the energy is gone, the incident is old news, and the items are overdue. The check exists to keep the loop honest.

And the pattern review, quarterly: the last several incidents, looked at together. AI incidents cluster. The same detection gap, the same unmonitored signal, the same eval blind spot — each incident individually explained, together they're a pattern with a name and a fix. The quarterly review is where the individual lessons become the system improvement.

The no-blame rule, made real

The AI incident post-mortem dies without the no-blame rule, and the rule has to be explicit because the natural instinct is to find the person. The operator who didn't notice. The engineer who shipped the model. The reviewer who approved the bad output. Blame feels productive. It isn't. It just guarantees the next incident gets hidden instead of reported.

The framing that works: the system failed, and the system is ours. The model, the monitoring, the thresholds, the review process, the training — all of it was designed by the team, and the incident is a finding about the design. The person who reports the incident first is the hero, not the villain. The teams that internalize this get early reports and fast lessons. The teams that don't get silence and repeat incidents.

The question that closes every post-mortem

End every post-mortem with the same question: if this exact incident happened again tomorrow, how would we find it faster?

The answer is the summary of the action items. If the answer is "the same way we found it this time," a customer complaint or a reviewer's hunch, the post-mortem failed, because the detection is still reactive. If the answer is "the override-rate signal would have caught it on day one," the post-mortem worked, because the next incident is cheaper before it even happens.

Think this argument fits your event? Tell me about the room — the calendar is selective.

Start a conversation