Skip to main content
ThunderLang
← All articles
ai-engineering

Test AI-Generated Ticket Creation for Duplicate Retry Side Effects

7 min read · 2026-10-10 · Allen Codewell

A large lens spans matching blank tags and an open compartment holding one card and one envelope.

AI-generated illustration. Accepting a retry requires checking both matching ticket IDs and the persisted ticket and notification-outbox rows.

Test AI-generated ticket creation by calling the real handler twice in an isolated fixture, then inspecting both responses and persisted effects. Returning the same ticket ID is necessary, but not enough. A retry must also avoid creating another ticket or notification-outbox row.

The acceptance question is concrete: what must you observe before accepting the agent’s change? This worked example turns a bounded requirement into assertions that catch a plausible implementation mistake.

The supplied excerpt from Hugging Face’s AutoSynthData: Generating Training Data for Enterprise Agents discusses turning environment-specific agent weaknesses into training tasks. It motivates the question, not the testing method below. This is an independent exercise, not a reported experimental result.

Declare exactly what a retry must preserve

Assume ticket creation writes to two database tables:

  • tickets stores the new ticket.
  • notification_outbox stores a pending notification for a separate delivery process.

The outbox row is a persisted effect. Creating it twice can leave two notification jobs even when the API returns the original ticket ID.

Use this bounded requirement:

After a successful ticket-creation request, a sequential retry with the same tenant, idempotency key, and identical payload must return the same ticket ID and leave exactly one ticket row and one notification-outbox row for that request.

Translate that requirement into an acceptance card:

Contract element Declared value or observation
Starting state Empty, isolated database; no delivery worker running
Tenant A
Idempotency key retry-17
Payload {"subject": "Printer offline"}
Execution order First call completes successfully before the second starts
Response condition Both calls return the same ticket ID
Ticket condition Exactly one row belongs to this tenant and key, with the declared subject
Outbox condition Exactly one row belongs to this tenant and key and references that ticket

This is narrower than “ticket creation is idempotent.” It names the inputs, ordering, and effects the test will observe.

Build a counterexample that returns the right answer

This illustrative Python handler looks up the original result in the ticket table. It reuses the stored ticket ID on retry, but mistakenly enqueues a notification on every call.

The fixture below uses only an in-memory database and never sends a notification. No execution result is claimed here.

def create_ticket(db, tenant, key, payload):
    row = db.execute(
        "SELECT id FROM tickets WHERE tenant = ? AND request_key = ?",
        (tenant, key),
    ).fetchone()

    if row:
        ticket_id = row[0]
    else:
        cursor = db.execute(
            "INSERT INTO tickets (tenant, request_key, subject) "
            "VALUES (?, ?, ?)",
            (tenant, key, payload["subject"]),
        )
        ticket_id = cursor.lastrowid

    # Deliberate bug: this also runs when returning an existing ticket.
    db.execute(
        "INSERT INTO notification_outbox "
        "(tenant, request_key, ticket_id) VALUES (?, ?, ?)",
        (tenant, key, ticket_id),
    )
    db.commit()
    return {"ticket_id": ticket_id}

A response-only test would miss the defect:

assert first["ticket_id"] == second["ticket_id"]

That checks result reuse, not side-effect reuse. The ticket uniqueness constraint in the next example cannot protect the outbox either: it constrains a different table.

This is intent drift in concrete terms. The requirement covers the returned result and stored effects; the implementation preserves only the result and the ticket row.

Test responses and committed state together

Place this acceptance test alongside the handler above. It creates fresh tables, makes two sequential calls, and queries the database after both calls have committed.

import sqlite3
from contextlib import closing


def test_successful_sequential_retry_has_one_effect_set():
    with closing(sqlite3.connect(":memory:")) as db:
        db.execute("PRAGMA foreign_keys = ON")
        db.executescript("""
            CREATE TABLE tickets (
                id INTEGER PRIMARY KEY,
                tenant TEXT NOT NULL,
                request_key TEXT NOT NULL,
                subject TEXT NOT NULL,
                UNIQUE (tenant, request_key)
            );

            CREATE TABLE notification_outbox (
                id INTEGER PRIMARY KEY,
                tenant TEXT NOT NULL,
                request_key TEXT NOT NULL,
                ticket_id INTEGER NOT NULL REFERENCES tickets(id)
            );
        """)

        tenant = "A"
        key = "retry-17"
        payload = {"subject": "Printer offline"}

        first = create_ticket(db, tenant, key, dict(payload))
        second = create_ticket(db, tenant, key, dict(payload))

        assert first["ticket_id"] == second["ticket_id"]

        tickets = db.execute(
            "SELECT id, subject FROM tickets "
            "WHERE tenant = ? AND request_key = ?",
            (tenant, key),
        ).fetchall()
        assert tickets == [(first["ticket_id"], "Printer offline")]

        notifications = db.execute(
            "SELECT ticket_id FROM notification_outbox "
            "WHERE tenant = ? AND request_key = ?",
            (tenant, key),
        ).fetchall()
        assert notifications == [(first["ticket_id"],)]

For the faulty handler, the expected failure is the final assertion: the outbox query would return two rows rather than one. Equal ticket IDs cannot compensate for that extra row.

The queries select every row for the tenant and request key before comparing the complete result. This matters. Filtering the outbox query by the expected ticket ID alone could hide an extra row referencing the wrong ticket.

For this narrow example, the minimal repair is to return immediately when an existing ticket is found. Replace the first branch with:

    if row:
        return {"ticket_id": row[0]}

The existing else branch still creates the first ticket, and only that path reaches the outbox insert. This repair addresses a completed request followed by a sequential retry, not every retry failure mode.

A mechanical dispenser holds one card while two envelope-shaped tiles collect beneath a separate chute.

AI-generated illustration. The illustrative bug reuses the ticket but inserts another outbox row on retry. A response-only assertion would miss that duplicate pending notification.

Apply the test to the agent’s actual change

The exercise calls its implementation directly, rather than substituting a canned response. Keep that boundary when adapting it: invoke the actual handler changed by the agent, backed by isolated test persistence. A passing test of this demonstration handler says nothing about unrelated application code.

Keep the assertions outside the implementation. A returned field such as deduplicated: true is not verification; inspect the stored effects.

Review test changes as carefully as handler changes. Removing the outbox assertion would weaken the contract, not fix the implementation.

Retain the requirement, fixture, assertions, code revision, and test result together. A passing result applies to the code and environment checked, not automatically to later revisions. It is regression evidence, not mathematical proof.

Decide which untested cases need another contract

The test’s limits are part of its meaning. These cases remain unresolved:

Untested case Separate decision to declare
Concurrent retries What must happen when two calls observe no existing result at the same time? Specify allowable responses and the final effect count.
Crashes between writes or around commit Must the ticket and outbox entry be atomic? How should a retry recover when the client cannot tell whether the first request committed?
Same key, changed payload Should the service reject the request, replay the original result, or apply another declared policy?
Same key, different tenant Should each tenant receive an independent ticket? Specify and test the isolation boundary.
Actual notification delivery What must the delivery worker do after retries or acknowledgements are lost? One outbox row does not establish one delivered notification.

The in-memory SQLite fixture also does not establish how another database behaves under concurrency or failure. If those semantics matter to the change, add isolated integration tests using the relevant database system.

Choose the next contract from the application’s retry model. If requests can overlap, prioritize concurrent retries. If processes can restart during creation, prioritize atomicity and recovery. Do not silently broaden what a passing sequential test means.

Acceptance checklist for the next AI-assisted diff

  • Retry scope: State the tenant, key, payload relationship, and whether calls are sequential or concurrent.
  • Observable effects: Check returned values and every persisted effect named in the requirement.
  • Fixture isolation: Use fresh test resources, no live customer data, and no external notification delivery.
  • Implementation boundary: Exercise the changed handler, not a mock that already returns the desired answer.
  • Counterexample: Check that the assertions reject a same-ID, duplicate-outbox implementation.
  • Untested failure modes: Record concurrency, crashes, payload changes, tenant isolation, and delivery as separate decisions where relevant.

The acceptance rule is simple: a correct response cannot compensate for an incorrect state change.

Use this bounded requirement to gate your first AI change with ThunderLang.