Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service

One morning a nightly run came back with 74 failures that my tooling had neatly filed under “stale credentials”. Seventy-four tests, one cause, one obvious fix: refresh some logins.

Except 52 of them were not stale credentials. The credentials were fine. A monthly operations task, the kind a human runs once per period, simply hadn’t been run yet, and the system said so in a very specific business error. Those 52 tests were failing for a reason that no amount of password-resetting would ever touch.

I only knew because I’d taught the test framework to read the one thing the test report didn’t show me: the response body. This article is about how that happened, and about the pipeline I built around it: deterministic rules first, an LLM only for the leftovers, and a set of sanity checks on everything the LLM says.

Why the report lies by omission

A huge share of API tests assert one thing: the status code. When that fails, the JUnit XML gives you something like:

Expected: 200
Received: 403

That’s all. Is it a permissions problem? An expired token? A business rule saying “not allowed right now”? The test report can’t tell you, because the test never looked. The answer was in the HTTP response, and the response was sitting in Playwright’s trace.zip the whole time.

So step zero of the test framework’s triage is trace-body enrichment: for each failure, read the actual response body out of the trace and append it to the failure text. Everything downstream, rules and LLM alike, sees the real error instead of a bare number. This single change is what split those 74 failures into 52 waiting on a monthly task and 22 that still looked like credential problems.

One practical note: response bodies can contain personal data, so they are cached in a bounded, in-memory FIFO (500 entries) and not written anywhere durable.

The pipeline

flowchart TD
    A["Failures, hundreds of them"] --> B["Enrich with the response body from the trace"]
    B --> C{"Deterministic rules"}
    C -->|match| D["Category. Done, free, repeatable"]
    C -->|no match| E["LLM categorizes the leftovers, in chunks of 25"]
    E --> F["Sanity checks on the model's answer"]
    F --> G["Remediation proposals. A human reviews, then applies"]

The idea is simple: most failures in a mature suite are old friends. The same five or six causes explain the bulk of any bad night. Paying an LLM to rediscover “the server returned a 500” on 200 tests is slow, costs tokens, and gives you slightly different category names each time. Rules do it instantly and identically.

Step 1: five rules, order matters

The rules are ordered, and the first match wins:

  1. Server or infrastructure error: a 5xx where a 2xx was expected.
  2. “Period not rolled”: one specific business error code, tied to that monthly ops task.
  3. Invalid payment method for this context: a known configuration mismatch.
  4. Stale credentials: a 401 or 403.
  5. Missing test account: a 404 on something that should exist.

Look at rules 2 and 4. The monthly-task error also satisfies the generic 401/403 “stale credentials” rule. If the generic rule ran first, I’d be back to my 74-failure fiction. The specific business-error rule sits above the generic 403 rule on purpose. Rule order is where your domain knowledge lives, so put the narrow, specific things before the broad ones.

The soft-assertion subtlety

Two details in the matching bit me, and both are worth stealing.

First, the rules anchor on the assertion library’s Expected: N / Received: N adjacency, not on “the text contains 403”. A status code can appear in a test title (“returns 403 for non-admin users”) that has nothing to do with what actually failed. Matching loosely means a passing expectation in the name of the test can masquerade as the cause.

Second, the rules check every Expected/Received pair in a failure, not just the first. The suite uses soft assertions, which keep going after a failed check and report all of them. A test might first fail a body-shape expectation (a missing field, say) and only afterwards fail the real status check. If you only read the first pair, you categorize by the symptom and miss the cause. Pseudocode:

for pair in all_expected_received_pairs(failure_text):
    for rule in ORDERED_RULES:
        if rule.matches(pair, enriched_body):
            return rule.category
return UNMATCHED   # goes to the LLM

Step 2: the LLM gets only what the rules couldn’t explain

Whatever survives the rules goes to the LLM, which groups it into root-cause categories. This is the part where a model genuinely earns its keep: weird timeouts, novel error text, things I haven’t seen before. And it’s the part where I stopped trusting it quickly, because models are great at judgement and surprisingly sloppy at bookkeeping.

Step 3: never let a model’s bookkeeping go unchecked

Categorization is a bookkeeping task dressed up as an intelligence task: here are N tests, put every one in exactly one bucket. Each failure mode below has its own defence.

A big single batch truncates silently. The model runs out of output and just stops, with no error. Fix: split into chunks of 25 and run them in parallel. Plus a coverage guard that notices truncation (the finish reason says “length”) instead of assuming the answer is complete.

The model identifies tests by positional index, and the indexes can collide with tests already handled by the rules. Fix: sanitize every index the model returns against the real input set. Anything it invented or reused gets dropped.

A test assigned to two categories. Fix: first assignment wins, deterministically.

A test assigned to none. Fix: anything unaccounted for lands in an explicit “Categorization Incomplete” bucket. The total in always equals the total out, so nothing silently vanishes. A visible “the model dropped these 6” is a feature; a quietly shorter list is a bug you find in a month.

Chunks named the same cause differently. Chunk one says “Timeout Errors”, chunk three says “Network Connectivity Issues”, same failures. Fix: a second, cheap call that sees category names only, never the test data, and merges duplicates. It’s a tiny prompt doing a tiny job.

assigned = {}
for idx, category in llm_output:
    if idx not in input_set: continue          # invented index
    if idx in assigned:      continue          # first wins
    assigned[idx] = category
leftover = input_set - assigned.keys()
assigned.update({i: "Categorization Incomplete" for i in leftover})
assert len(assigned) == len(input_set)         # the invariant

That last assert is the whole philosophy: the model proposes, plain code checks the arithmetic.

Step 4: remediation, with a human in the loop

Categorizing a failure is half the job. The test framework then tries to say what to do: it maps the failure to the scenario’s data file and decides whether it’s fixable. Typical fixes:

  • backfill missing metadata in a scenario,
  • swap an out-of-stock product for a candidate from a resolver pool, after probing that the candidate really is available,
  • diagnose a stale static user: verify the login, reset it, or reuse another.

Every proposal carries a “certain / not certain” flag, a reason, and a “requires human review” marker. Applying a fix can immediately re-run the one test to give a per-test verdict, and the test framework also prints a ready-to-paste command so a person can verify independently without trusting the test framework’s word.

Honest limit: it can’t fix everything. Scenarios built by a payload-builder function can be regenerated; roughly 350 legacy scenarios can’t, and the test framework reports those as “not fixable” instead of pretending.

Persistent failures: telling “broken” from “flaky”

A single bad night tells you little. The test framework also looks back over the last 30 days of daily results, in batches of 5, and computes per test:

  • first and last failure,
  • consecutive failing days,
  • flakiness, as failed days divided by total days,
  • a “flaky” label when a test both failed and passed on different days.

A test failing nine days in a row is a different conversation from one that fails every third day. The first needs an owner; the second needs a look at timing or shared data.

One subtle rule: a day whose results can’t be fetched is skipped. It is not counted as a pass, and not as a fail. Treating “I couldn’t look” as “it passed” quietly deflates flakiness; treating it as “it failed” inflates it. Unknown stays unknown.

The denominator lesson

I’ll end with a humbler mistake that is the same bug as the credentials story in a different costume.

While working on the server, I ran a type-check. It came back clean. Great news, except it had checked zero files. There was no TypeScript config in that directory, so the tool had nothing to examine, and “nothing examined, nothing wrong” looks exactly like “everything examined, nothing wrong”.

That’s the rule I now apply everywhere, and I call it checking the denominator: a green result means nothing until you know how many things were examined. Zero errors out of zero files. Zero failures because the runner found zero tests. 74 “stale credentials” that were really one business error nobody looked at.

The triage pipeline bakes this in: the “Categorization Incomplete” bucket, the input-set assertion, the skipped-day rule, the “unknown, not 0” handling. Every stage states what it examined and how much it managed to account for.

Adding the next known failure

When a new recurring pattern shows up, adding it is a recipe, not a project:

  1. Add a rule above any broader rule that overlaps it.
  2. Add the matching remediation branch.
  3. Update the tool descriptions so the agent (more on that in article 4) knows it exists.
  4. Verify end-to-end on a real run, not on a hand-built fixture.

Every pattern you promote from “ask the LLM” to “rule” makes the next night cheaper, faster, and more consistent. The LLM’s job shrinks toward the genuinely strange, which is exactly where you want it.

Takeaways: enrich failures with the evidence the report omits; let cheap deterministic rules take the bulk and put specific rules before general ones; use the LLM for judgement, then check its bookkeeping with plain code; and before you believe any green, ask how many things it actually looked at.

Next: Article 3, Comparing responses across a flag: how to prove a migration behaves identically with a feature flag on and off.