Spec Replay Regression: What Happens When an Agent Rewrites a Decision Already Made
The Problem Has a Name: Spec Replay Regression
Spec replay regression happens when an AI coding agent regenerates a solution your team already evaluated and deliberately rejected. The agent produces code that compiles, passes tests, and satisfies the current spec text, but silently reintroduces a behavior, architecture, or data handling approach that was explicitly ruled out earlier in the project's history.
The key word is silently. Nothing in the build pipeline turns red. The agent wasn't being careless. It simply had no access to the rejected path. From its perspective, it was choosing among valid options. From the team's perspective, a settled decision just got undone.
Why This Happens: Agents Reason From State, Not History
Most AI coding agents receive a spec or task description that describes what the system should do right now. That is a snapshot of current intent. It does not carry the reasoning history of how the team arrived at that snapshot, and it does not carry the list of things the team considered and said no to.
Human engineers absorb rejection history informally. A senior developer who was in the room when the team ruled out client-side session tokens for compliance reasons will not propose them again six months later. An agent starting a new context window has no such memory. It reasons from the spec it sees, identifies a technically valid approach, and implements it. If the spec doesn't say "do not use client-side session tokens," the agent has no reason to avoid them.
This is not a hallucination problem. The agent isn't inventing something false. It is doing exactly what it was designed to do: find a conformant implementation. The failure is that conformance was defined only in the positive direction.
A Worked Example: The Session Token Case
Assume a simplified spec (expressed here as structured pseudocode for clarity):
SPEC AuthFlow v1.4
REQUIRE: user session persists across page reloads
REQUIRE: session invalidated after 30 minutes of inactivity
REQUIRE: session data not accessible to client-side JavaScript
An agent generating an implementation for this spec might choose between several approaches:
- HttpOnly cookies backed by server-side session store
- localStorage tokens with a client-side expiry check
- sessionStorage with a heartbeat mechanism
Approaches 2 and 3 violate the third requirement as written, so a conformance check would catch them. Good.
Now suppose three sprints earlier, the team also evaluated and rejected a fourth approach: signed JWTs stored in HttpOnly cookies, where the JWT payload encodes authorization claims directly. The team rejected this because it creates a window between token issuance and revocation where a compromised token cannot be invalidated before expiry. That decision was captured in a design review document and a Slack thread, not in the spec.
The current spec does not forbid JWTs. A new agent, or the same agent in a new context, generating a session implementation might reasonably choose signed JWTs in HttpOnly cookies. The implementation would satisfy all three written requirements. A positive-only conformance check would pass it.
The ghost of the rejected option came back, and nothing caught it.
The Gap: Positive Conformance Is Not Enough
A verify-diff that checks whether agent output satisfies declared requirements answers one question: Does the code do what the spec says it must do? Necessary. Not sufficient.
The missing question is: Does the code do anything the spec said it must not do?
These are different logical operations. Positive conformance checks membership in an allowed set. Rejection checks verify exclusion from a prohibited set. If your spec only encodes the allowed set, rejection checking is structurally impossible, regardless of how sophisticated the verification tooling is.
The spec itself has to change. Rejection decisions must become first-class spec citizens, not footnotes in meeting notes.
Making Rejection Decisions Machine-Verifiable
Durable proof artifacts need a richer vocabulary. A spec that can only say REQUIRE and ALLOW cannot capture the full intent of a team that has made deliberate exclusion decisions.
Consider extending the session token spec with explicit rejection annotations:
SPEC AuthFlow v1.4
REQUIRE: user session persists across page reloads
REQUIRE: session invalidated after 30 minutes of inactivity
REQUIRE: session data not accessible to client-side JavaScript
REJECT: localStorage or sessionStorage as session backing store
REASON: violates client-side inaccessibility requirement
DECIDED: 2024-03-12, sprint-14 design review
REJECT: self-contained signed tokens (e.g. JWT) as session mechanism
REASON: no server-side revocation before expiry; compliance risk
DECIDED: 2024-03-12, sprint-14 design review
REFERENCE: ADR-0041
Now the spec carries both sides of intent. A verify-diff running against agent output can check not only that the implementation satisfies REQUIRE clauses but that it does not exhibit patterns corresponding to REJECT clauses. The rejection is machine-readable, timestamped, and linked to its rationale.
When the agent produces a JWT-based implementation in sprint 22, the verify-diff catches it. Not because the agent made a logical error, but because the spec now knows what the team said no to.
What "Durable" Means for Rejection History
Proof artifacts become stale when the spec evolves but the artifacts don't. Rejection history faces a particular staleness risk: a REJECT clause that made sense in sprint 14 might be revisited in sprint 30 when a new compliance framework changes the risk calculus. Treating REJECT clauses as immutable historical records turns the spec into a graveyard of outdated prohibitions.
A useful distinction is between closed rejections and conditional rejections.
A closed rejection says: this approach violates a hard constraint (legal, security, platform) that is unlikely to change. Verify against it indefinitely.
A conditional rejection says: this approach was ruled out given current constraints X and Y. If X or Y change, the rejection should be reviewed. The proof artifact should flag the relevant conditions and surface the rejection for human review when those conditions are revisited.
This doesn't make verification automatic. It makes the conversation explicit. Instead of an agent silently undoing a past decision, the verify-diff produces a finding: "This implementation pattern matches a conditional rejection tied to compliance context ADR-0041. That context was last reviewed in sprint 14. Review before accepting."
The agent doesn't make the call. The team does. But the team now has the information.
The Verify-Real-Code Loop and Rejection Awareness
When an agent checks its own diff against intent before handing work back, that loop has the same scope limitation as any other conformance check if the spec only encodes positive requirements. The agent can verify "I satisfied all REQUIRE clauses" and still be blind to "I implemented a pattern the team explicitly rejected."
Rejection-aware specs close this loop. The agent, checking its own diff, now has access to the full intent surface: what the system must do, what it must not do, and why. It can surface a conflict before any human review step. If the agent's reasoning led it to a rejected pattern, it can backtrack and choose a different path, or flag the conflict for human decision rather than presenting a confident, passing implementation.
That changes the agent from a generator of conformant code into a participant in settled decisions. The difference is meaningful when the cost of replaying a rejected architectural decision is measured in compliance risk or security exposure rather than a failed test.
An Exercise: Audit Your Spec for Ghost Options
If you work with AI-generated code and spec-driven development, try this:
- Pick one component that has been through at least two design cycles.
- Retrieve all design review notes, ADRs, and discussion threads for that component.
- List every approach that was explicitly evaluated and rejected.
- Check how many of those rejections appear anywhere in the current spec.
In many codebases, the answer to step 4 is zero or close to it. The spec records what survived selection, not what was considered and discarded. Every absent rejection is a ghost option that an agent operating on that spec could legitimately implement.
The exercise works because it makes the incompleteness concrete. Once you can see which decisions are undocumented in the spec, you can prioritize which to encode first. Start with rejections that carry the highest risk if replayed: security and compliance decisions are the natural candidates.
Spec Memory Is a Team Property, Not an Agent Property
It is tempting to frame this as a model capability problem: if only the agent had a longer context window or better retrieval, it would remember the rejected paths. That framing is misleading. Even an agent with access to every document your team has ever written would need those documents to explicitly record rejection decisions in a machine-readable, verifiable form. Burying a rejection in prose in a Slack export is not the same as encoding it as a checkable constraint.
Spec memory has to live in the spec.
The spec has to be durable enough to persist across agent context windows, team member turnover, and the passage of time. The proof artifacts generated from that spec have to stay fresh as both the spec and the codebase evolve. That is a tooling and process problem, not a model capability problem. It is approachable with existing methods, if the spec format is designed to carry the full intent surface rather than just the positive requirements.