Skip to main content
ThunderLang
← All articles
ai-engineering

Your Spec Should Be the Runbook: Intent Verification at Incident Time

7 min read · 2026-09-28 · Allen Codewell

The Problem With Runbooks When AI Writes the Code

A runbook is only as trustworthy as the person who wrote it. When a human engineer wrote the service, the runbook could reasonably summarize their intent: restart this process, check that queue depth, roll back to this tag. The author of the code and the author of the runbook were close enough in time and context that the two documents roughly agreed.

That assumption breaks when an AI agent authors the code. The agent does not write a runbook. It produces a diff, merges it, and moves on. If an incident occurs six weeks later, the on-call engineer inherits behavior they did not design, in a codebase they may not have reviewed line-by-line, with no canonical document explaining what the code was supposed to do versus what it happens to do under pressure.

This is not a criticism of AI-assisted development. It is a structural observation: agent authorship severs the implicit link between the code's behavior and any human's memory of its intent. The runbook gap is not a process failure. It is an architectural one.

Intent as the Only Durable Record

Observed behavior is not the same as declared intent. A service might be passing all its health checks and still be operating outside the boundaries its designers specified. Conversely, an alert might fire because observed behavior deviates from a threshold that was itself never grounded in any requirement.

When an AI agent handles incident response, it is optimizing for observed signals: error rates, latency percentiles, crash loops. It has no intrinsic access to the original intent behind the system unless that intent is captured somewhere machine-readable. Without that record, automated remediation is pattern-matching against symptoms, not reasoning about what the system was designed to guarantee.

A machine-verifiable spec changes this. If the spec declares that a payment processing endpoint must reject duplicate transaction IDs within a 60-second window, that constraint exists independently of any implementation. An agent responding to a payment incident can ask: does the proposed remediation preserve that constraint? Without the spec, the agent can only ask: does the proposed remediation reduce the error rate? Those are different questions, and under incident pressure, the second one often produces answers that satisfy the alert while violating the design.

What a Spec-as-Runbook Actually Looks Like

Treating a spec as a runbook is not about writing prose documentation that happens to live next to the code. It means the spec must be:

Machine-readable and checkable. A human-readable requirements document cannot be verified against a diff automatically. The spec needs enough formal structure that a tool can compare a proposed code change against it and produce a conformance result.

Intent-bearing, not behavior-describing. A spec that says "the retry count is 3" describes current configuration. A spec that says "retries must not exceed the point at which downstream services are at risk of cascade failure" captures intent. The first becomes stale the moment someone changes a constant. The second remains valid as the implementation evolves.

Versioned alongside the code, but independently authoritative. The spec should track what was intended at each point in the system's history, not just what was deployed. When a postmortem asks "was this behavior intentional?", the spec version at the time of the incident should be able to answer that question.

Executable at incident time. The spec does no good sitting in a wiki. At incident time, the on-call engineer or the responding agent needs to be able to run a verification check against a proposed remediation and get a pass/fail result, not read a document and make a judgment call.

A Worked Example: The Remediation Verification Loop

Assume the following setup (this is a hypothetical example, not a measured result):

A team runs an order fulfillment service. Their spec declares two constraints relevant to this example:

  1. Orders in a pending state must not be transitioned directly to cancelled without first attempting a retry if the failure reason is a transient network error.
  2. Inventory reservation must be released atomically with any order cancellation.

An AI agent detects a spike in failed order transitions and proposes a remediation: add a background job that scans for stuck pending orders older than 10 minutes and marks them cancelled.

Without a spec-as-runbook, the on-call engineer reviews this diff under pressure. The diff looks reasonable. The background job would clear the queue and stop the alerts. They apply it.

With a spec-as-runbook, the verification step catches two problems before the remediation is applied:

  • The proposed job does not inspect failure reason before cancelling, so it would cancel orders that should be retried, violating constraint 1.
  • The job marks orders cancelled in a loop without a transaction wrapping the inventory release, creating a window where the order is cancelled but inventory is not freed, violating constraint 2.

The value here is not that the spec is smarter than the agent. The value is that the spec makes the intent of the system legible at the moment when the pressure to act fast is highest. The agent's proposed fix was locally reasonable and globally wrong. The spec provides the standard against which "globally wrong" can be determined automatically.

Why This Changes the Incident Postmortem

Postmortems in traditionally authored systems typically ask: what did the system do, and why did a human let that happen? The postmortem is fundamentally a narrative about human decision-making.

Postmortems in agent-authored systems need a different frame. No human decided the code would behave a particular way. The question shifts to: what did the system do, does that behavior conform to the declared spec, and if not, at what point did spec and implementation diverge?

If the spec is the runbook, this question is answerable. You can produce a diff between the spec at time T and the implementation at time T and determine whether the incident was caused by a spec violation, an implementation drift, or a gap in the spec itself. That is a defensible, auditable recovery story. It is also the only kind of postmortem that makes sense in a world where the code's author may have been an agent that no longer exists in its original form.

Teams that lack this artifact are left reconstructing intent from commit messages, Slack threads, and memory. Under governance pressure, that reconstruction is fragile. Under legal or regulatory scrutiny, it may be inadequate.

The Spec Disciplines the Agent, Not Just the Code

There is a subtler benefit worth naming. When agents know they will be checked against a spec before their changes are accepted, the spec becomes a constraint on how the agent generates code in the first place. An agent operating in a verify-diff loop does not just produce code and wait for a human to review it. It can check its own output against the spec before proposing a change, identify constraint violations, and revise.

This closes the loop between intent declaration and code generation in a way that post-hoc review cannot. Post-hoc review catches violations after the fact. Spec-gated generation prevents a category of violations from being proposed at all.

The incident-time value of the spec is therefore partly a function of how rigorously the spec was used during development. A spec that was only consulted at postmortem time is thin. A spec that constrained every agent-generated change is rich, versioned, and reflective of every deliberate decision the team made about intended behavior.

What Remains Uncertain

This approach rests on an assumption that needs to be stated clearly: it requires that specs be written with enough precision to be machine-verifiable, and that they be kept current as the system evolves. Both of these are non-trivial. Specs that drift from the codebase are worse than no specs, because they produce false confidence at incident time.

The discipline of keeping a machine-verifiable spec fresh as an agent-driven codebase changes is a genuine ongoing cost. That cost is real, and teams should weigh it honestly. The argument here is not that spec-as-runbook is free, but that the alternative, reconstructing intent under incident pressure from a codebase you did not write, is more expensive and less reliable.

A Practical Exercise

Pick one service your team owns where AI has contributed a meaningful portion of the recent commits. Without looking at any documentation, write down the three behavioral guarantees you believe the service makes to its consumers. Be specific enough that a tool could check a diff against them.

Now look at the codebase and ask: is each guarantee verifiably encoded anywhere, or does it live only in your memory and the memories of whoever was in the room when the design decisions were made?

If the guarantees live only in memory, you already have a runbook gap. The question is whether you discover it during a quiet retrospective or during an incident at 2 AM with an agent-generated remediation waiting for your approval.


If you want to start closing that gap, gate your first AI change against a declared spec with ThunderLang's verify-diff workflow.