Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service

Here is a sentence I have learned to distrust: “The suite is green with the flag on, so the migration is safe.”

It is not. Green means every assertion the test author thought to write held up. Most API tests assert a status code and maybe two fields. Meanwhile the response has sixty fields, and one of them is a price that is now off by a cent, or a date that quietly lost its timezone. The test never looked there, so the test never complained.

A migration has a stricter promise than “the tests pass”. The promise is: the new path behaves identically to the old one. To check that, you do not need better assertions. You need to put the two responses next to each other.

That is what the response comparison tool in my test framework does.

The axis is a flag, not a version

The usual framing for this problem is “compare v1 to v2”. Mine is slightly different. We migrate behind a feature flag: with the flag off, a request is served by the old data source; with it on, by the new one. Same endpoint, same URL, same request. Only the flag changes.

So the thing I vary is a flag value. (The same machinery would work for comparing two API versions. Nothing in it cares what the axis is, only that there are labelled runs to line up.)

The flow

flowchart TD
    P["Create a project"] --> OFF["Run the suite with the flag off"]
    P --> ON["Run the suite with the flag on"]
    OFF --> CAP1["Capture request and response pairs"]
    ON --> CAP2["Capture request and response pairs"]
    CAP1 --> CMP["Compare every pair of runs"]
    CAP2 --> CMP
    CMP --> STR["Structure pass: same shape?"]
    STR --> VAL["Value pass: shared paths only"]
    VAL --> R["Report"]

Concretely:

  1. I create a project, a named container for one comparison.
  2. I run the same Playwright suite once per flag value. The flag is forced inside the tests, so each run is a clean sample of one world.
  3. The test framework opens each run’s Playwright traces and pulls out every request/response pair: the response body, the request id, the duration, and a ready-made cURL.
  4. It compares the runs.

Nothing here needs a special test suite. I reuse the one I already have. The tests are just a convenient way to generate realistic traffic, and the traces already record everything that crossed the wire.

Storage is simple. For each flag value and each endpoint, the test framework keeps the response JSON, a “structure” JSON (more on that in a second) and a bit of metadata. Re-capturing a flag value wipes that value’s prefix first, so stale data cannot leak into a new run, and endpoints that no longer appear get pruned. You need at least two captured runs, and the tool compares every pair, so adding a third flag value does not mean rewriting anything.

Pass one: structure

The first question is cheap: does the response have the same shape?

The trick is to throw the values away. I take each response and:

  • replace every leaf value with null,
  • sort the keys,
  • keep only the first element of every array (a list of 40 items has the same shape as a list of 1),
  • flatten what is left into dotted paths.

So a response like this:

{
  "id": "a1",
  "items": [
    { "sku": "X1", "price": 12.1, "tax": 1.0 },
    { "sku": "X2", "price": 8.0,  "tax": 0.6 }
  ],
  "shipping": { "method": "ground" }
}

becomes the paths:

id
items[].sku
items[].price
items[].tax
shipping.method

Now comparing two responses is set arithmetic. Anything that appears in only one run’s path list is a structural difference, reported as “only on this side”. No judgement, no fuzziness. It catches the loud failures: a field that disappeared, a field that is new, an object that became a string.

Pass two: values

Structure matching proves the shapes agree. It says nothing about whether the numbers do. A response can have a perfect items[].price path and the wrong price. Right shape, wrong value is the migration bug that sails past every status-code assertion.

So the second pass walks the two responses leaf by leaf and compares values. Three rules keep it honest instead of noisy:

Only shared paths are compared. If items[].tax exists only in one run, the structure pass already reported it. Reporting it again as a value difference would double-count the same problem and bury the report.

null equals undefined. One run says "note": null, the other omits note entirely. In my captures this is almost always serialisation noise, not a real difference, and flagging it would drown the real findings. Numbers within 0.005 are equal. Floating point is a small liar. One side returns 12.1 and the other 12.099999999999998. Nobody in finance would call those different, and neither does the tool.

A tiny worked example

The same endpoint, captured twice:

// flag OFF
{ "id": "a1",
  "items": [ { "sku": "X1", "price": 12.1,               "tax": 1.0 } ],
  "discount": null }
 
// flag ON
{ "id": "a1",
  "items": [ { "sku": "X1", "price": 12.099999999999998, "tax": 1.2 } ],
  "currency": "USD" }

Structure pass: currency exists only in the ON run. (discount is null in OFF and absent in ON, but at the structure level that is a path that exists on one side only, so it shows up here too. It is exactly the kind of thing worth a human glance.)

Value pass, shared paths only:

pathOFFONverdict
ida1a1equal
items[].skuX1X1equal
items[].price12.112.099999999999998equal (within 0.005)
items[].tax1.01.2different

The price noise is absorbed. The tax difference is the real finding, the sort of thing every test with “status is 200” would have waved through. That is the whole point of the tool in one table.

Each endpoint then gets a verdict: MATCH, MISMATCH, or MISSING (it exists in one run and not the other). The value status is tracked separately from the structure status, because “same shape, different numbers” and “different shape” are different conversations with different people.

Making “the same endpoint” mean the same endpoint

Here is a problem you only meet when you actually run this. The OFF run created an account with one id and the ON run created a different account with a different id. So the paths differ:

/accounts/8f14e45f-ceea-467a-9575-1b2f0e0c3a11/documents/1042
/accounts/3c59dc04-8e88-4d1b-b1d4-6d2a77c5e7aa/documents/2977

They are the same logical endpoint. A naive comparison would treat them as two unrelated endpoints and report “missing” on both sides.

So the test framework normalises paths before lining anything up. Hashes, UUIDs and numeric ids in paths collapse into placeholders:

/accounts/{id}/documents/{id}

An endpoint’s identity is then test name + method + normalised path. Including the test name matters: the same endpoint called by two different tests with different intent stays two separate comparisons, instead of one test’s response being compared with the other’s.

On a sample project this lined up roughly 445 endpoints. I would not want to do that matching by hand, and I would not trust myself to do it consistently by eye.

The performance comparison costs nothing

While capturing each pair I also record how long it took. That means once the data is there, I get a performance comparison across flag values for free: same endpoint, same request, new data source versus old one.

“The new path is correct but noticeably slower” is an extremely useful thing to know before the flag goes to everyone, and it needs no performance-testing code. It falls out of data you are already holding. This is a pattern I keep rediscovering: if you capture rich data once, new questions become cheap.

(It is not a load test, to be clear. One request per endpoint tells you about relative latency, not about behaviour under pressure. For that, there is a different tool, which is the subject of a later article.)

Recovering test titles, the unglamorous bit

A small problem that took longer than it deserved: after a run, Playwright’s output folders are named with a hash suffix, not a readable test title. If the report says “mismatch in a1b2c3-retry1”, nobody learns anything.

The title is, however, inside the trace. The first line of the trace file in the zip contains it. So after capture the test framework reads that line and recovers the human-readable title for each pair. It is a few lines of code, and it is the difference between a report people read and a report people close.

What comes out

A markdown report you can paste into a ticket or a review, and the same thing as JSON for anything that wants to consume it. The test framework keeps the last 20 comparisons, and projects can be downloaded and uploaded so a colleague can open the exact data I was looking at instead of re-running everything.

Why the MCP tools here are thin

The test framework also exposes this to an AI agent through MCP (Model Context Protocol, the standard way to give an agent callable tools). Archives and comparisons have a small set of tools, and they are deliberately thin: they expose projects and reports and do not re-implement trace parsing.

That was a conscious decision. The UI already knows how to compare. The agent does not need a second, slightly different copy of that logic. When it needs to dig into a specific mismatch, it can read the raw trace itself, the same file I would open, and reason about it. Duplicating the parsing in a tool would give me two implementations to keep in sync and a new way for the UI and the agent to disagree.

The rule I follow: put the deterministic, repeatable counting in the test framework (the comparison itself), and leave the open-ended “why did this happen?” to the agent, which is good at exactly that and has the raw material to do it.

Takeaway

If a change has to be invisible to users, test for invisibility directly. Capture both worlds, line them up, and compare. Structure first, because it is cheap and loud. Values second, with enough tolerance that float noise does not cry wolf. Normalise identities so you compare like with like.

And keep the suite you already have. It was already generating the traffic; I just started keeping the receipts.

Next: 4. Making the test framework agent-native, where I explain why an analysis you ask an AI for every day should become a tool.