Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service

I built the test framework for QA. Then I noticed that most of what it shows matters to someone who has never opened a test file: a security-alert count matters to a lead, a stale readiness checklist to a PM, a failing repro command to a developer.

So the test framework is a website the whole team can open, and each person walks through a different door. This article is a tour organised by who walks through which door, followed by the argument I think matters most: why I built these instead of buying them, and what that honestly costs.

The shape of the site

The test framework is a React app with a sidebar. Each tool is a section. Providers are hoisted above the sections, so a filter you set on one page survives a trip to another and back. There is a dark mode that follows the system by default. These are small things, and they are why people kept coming back instead of bookmarking one page.

Here is the audience map I ended up with:

block-beta
    columns 1
    H["Test framework"]
    QA["QA — triage, bug tracker, coverage, runner"]
    DEV["Developers — failure report, traces, repro command, comparison"]
    LEAD["Engineering leads — trends, flakiness, alerts, load tests"]
    PM["PMs and support — readiness queue, tickets, meeting notes"]

Articles 2 to 4 covered the triage and comparison tools. This one covers the rest.

For QA: the bug tracker that notices when a tag goes stale

Our tests carry a @bug tag when a test is known to fail because of a real defect. Ticket gets filed, tag goes on, the test stays red on purpose so nobody forgets.

The problem is that two systems now hold the same fact: the tag in the code and the status in the ticketing system. They drift. A developer fixes the bug, the ticket closes, and the tag stays for months. The test is green and still labelled as a known bug.

The bug tracker reconciles three sources: the tags in the test tree, the tickets, and the latest results. That gives four useful buckets:

  • Still failing: tagged, ticket open, test red. Working as intended.
  • Now passing: tagged, test green. Probably fixed. Worth a look.
  • Untracked failures: red test with no tag and no ticket. Somebody should own this.
  • Stale tags: the ticket is resolved but the tag is still there. This is the one nobody finds without a tool, because nothing is visibly broken.

The stale-tag idea is almost embarrassingly simple. It is a join across three lists. It is also exactly the thing that never gets done by hand, because the manual version is “open every tagged test and check its ticket”. Nobody has that afternoon.

For QA and developers: the repro loop

Developers mostly want one thing from a failed run: what request went out, what came back, and how do I make it happen again. The failure report gives the first failed assertion per test in a sortable table (matcher, expected, received, HTTP code, message) with CSV export. The trace viewer shows the request and response pairs. And there is a copyable command that re-runs exactly that test.

None of this is novel. What matters is that it is one click away from the failure instead of a conversation that starts with “can you send me the logs”. A developer who can paste a command and see the failure on their own machine will fix it faster than one who is reading a screenshot.

The landing page of the test framework is test performance, because that is the question leads ask first: are we getting better or worse?

It reads the nightly runs and shows:

  • duration, pass rate and flakiness over time
  • a 7-day strip for the quick glance
  • a failure calendar, so a bad Tuesday is visible as a bad Tuesday
  • a persistent-failures panel: tests failing N days in a row
  • a base-run versus compare-run view, with an order-tests versus API-tests filter
  • a button that copies the command to re-run every test that changed status

Two design decisions are worth stealing.

Store the breakdown, not just the total. Per-day results are stored with their breakdown, so sparklines can be re-scoped (all tests, only one family) without refetching anything.

Say “unknown”, not zero. Older data sometimes lacks a field. A chart that shows 0 for “we did not record this” is lying politely. The test framework returns “unknown” and renders it as such. I will take a gap in a chart over an invented number every time.

Flakiness is defined plainly: failed days divided by total days, and a test counts as flaky only if it both failed and passed on different days. A test that fails every day is not flaky. It is broken, and it lives in the persistent-failures panel instead. A day whose results cannot be fetched is skipped, never counted as pass or fail. Our rough goal is flakiness under 5%, with the pass rate sitting in the low-to-mid 90s. Those two numbers on one page make for a short, honest status meeting.

Dependency alerts, and the rate-limit floor

Leads also own the question “what security debt are we carrying?“. The dependency tracker shows open security alerts plus the automated update pull requests, with a “new in 24h” marker, a daily history reconstructed from an event log, and a count pill in the sidebar. The pill is there so the number is visible from every page, not only this one.

Two details I like:

  • Plain updates are shown as a neutral “update”. They never get a made-up severity. An update PR is not a vulnerability, and pretending otherwise inflates the numbers and trains people to ignore the page.
  • It was the first background refresh loop in the app, and that is where I had to be careful. The loop refreshes every 15 minutes. There is also a hard 60-second floor between live calls to the source host, enforced even when someone presses the refresh button ten times in a row. If the source answers with a rate-limit response, the loop backs off. If a call fails, the page serves the last good cache instead of an error.
flowchart TD
    R["Refresh requested"] --> Q{"Last live call under 60 seconds ago?"}
    Q -->|yes| C["Serve cache"]
    Q -->|no| S{"Source says slow down?"}
    S -->|yes| B["Back off and serve cache"]
    S -->|no| F{"Call failed?"}
    F -->|yes| L["Serve the last good cache"]
    F -->|no| LIVE["Live call, then update the cache"]

The floor is the important line. A dashboard that hammers its data source because a human got impatient is a dashboard that gets you rate-limited for everyone, including the automation that needs the same source.

Load testing with custom k6 reports

Load testing is where leads and QA meet. I use k6, the open-source load tool, and the test framework wraps it so the whole team can launch and read runs without touching a terminal.

A scenarios ladder. You do not start with the big test. The ladder is: a one-virtual-user happy path (does the journey work at all), a smoke test, a baseline, then a burst. Each rung is a launch button. If rung one fails, nobody wastes an hour on rung four.

Flows as reusable journeys. A scenario says how much load; a flow says what each virtual user does. We have a full checkout, a browse, an admin flow, mixed traffic, and an “exception flood” that deliberately sends bad requests. The same flow runs under any rung.

A report that answers our questions. k6’s own terminal summary is fine for one person at one keyboard. The test framework’s report has tabs for requests, endpoints, timing, HTTP status, checks, data and logs, plus orders created and notes per run. Two features exist to remove a round trip:

  • Failed requests come with a cURL. When a request fails under load, the first thing a developer does is try to reproduce it. Having the exact command next to the failure removes a round trip.
  • Recovery from logs. Long runs get interrupted: the server restarts, someone closes a tab. Instead of losing the run, the test framework can rebuild it from the logs. The server also parses k6’s terminal summary as a fallback when the structured output is missing. Boring plumbing, but it means a run is never just gone.

I also track a proxy metric for “out of stock” style 4xx responses on the cart, quote and order endpoints, and a separate count of 5xx. Under load, a 4xx that means “the test ran out of data” and a 5xx that means “the system fell over” are very different stories, and a single error rate hides that. Thresholds and targets live in one config, so the report and the launcher never disagree about what “passing” means.

For scale, I will give two data points and no more. A real event handled roughly 1,600 orders in five minutes in production. An earlier test found that 80 to 120 concurrent account-creation calls failed. One real event, one earlier test: they are anecdotes with a shape, not a benchmark suite, and I would not generalise from either.

For PMs and support: reading tickets without reading code

Not everyone who needs to understand a ticket reads diffs. The ticket and PR reader fetches a ticket and its linked pull requests as clean markdown, with a “copy for AI” button and a grounded Q&A chat over the ticket and PR. The context is capped, so a giant diff cannot blow up cost or quietly truncate.

The readiness queue and its token governor

On top of the reader sits the QA-readiness queue. A background job walks the sprint’s tickets, reads each one with its PR diff, and scores it against a checklist:

  • is the feature flag named?
  • is there a working cURL?
  • does it use the right kind of test token?
  • are there examples for more than one market?
  • are there log links?

It uses a small, cheap model, because this is checklist work, not deep reasoning. Results are cached along with the model and token usage, so the page shows what was spent as well as what was found.

A job that reads dozens of diffs in a loop will hit provider rate limits unless something manages it. So it has a token-per-minute budget governor. It watches the rate-limit headers on responses, and when the budget is nearly spent, it pauses and the page shows “waiting until…“. The UI never looks hung, and the job never gets throttled into a wall of errors. The governor is simple arithmetic, and it is the difference between a feature and an incident.

The payoff is meant to be social, not technical. A ticket that reaches QA without a flag name or a usable cURL costs a round of questions; with the queue, the author can see the gap before the handoff.

Meeting notes that name their tickets

Anyone who has left a planning meeting with “we decided stuff about the thing” will see the point. You paste a transcript, and the test framework:

  1. Extracts ticket mentions. The prompt deliberately over-includes, because a false positive costs one click and a missed ticket costs a lost decision. It copes with speech-to-text quirks such as digits spoken aloud.
  2. Lets you discard the false positives.
  3. Generates a note per ticket from a focused window of transcript lines: 5 before each mention, 10 after. Feeding the whole transcript to every ticket would be expensive and would blur which discussion belonged to which ticket.
  4. Publishes each note as a comment on its ticket.

When the extractor is unsure which ticket was meant, the suggestions come from plain keyword search with stopwords. No model call. A keyword match is deterministic, free and good enough for “did you mean one of these”.

Context pages: where the system prompts live

One small feature quietly serves everyone. The test framework has markdown “context documents”: an overview, conventions, backend domain knowledge, a ticket-writing format. They are literally the system prompts behind the AI features, and they are viewable and editable in the UI. A PM can fix a wrong sentence in the domain doc without opening an editor or a repo. Roadmaps work the same way. Prompts that humans can read and correct beat prompts buried in code that only I can change.

Buy versus build

Everything above could in principle be bought. Test-management suites, dependency dashboards, load-test clouds and meeting summarisers all exist as subscriptions. So why build?

A subscription gets you someone else’s idea of a dashboard. It answers the questions the vendor’s average customer asks. Mine were: which @bug tags are stale? Which failures share one root cause? What is our order-test flakiness against API-test flakiness? No product ships that, because it is shaped like our data and our conventions.

Building got cheap. The scarce skill in testing is no longer writing tests or even writing UI. It is knowing which question to ask. With an AI assistant, a new page over data I already had was a matter of hours, not a quarter. The cost curve moved, and the decision should move with it.

It owns its data model. Because the test framework reads our own results manifest, tags, and tickets, joins that are impossible across three separate products (tag against ticket against last night’s result) are a few lines of code.

Now the honest part, because build-versus-buy posts that skip it are advertising:

  • You maintain it. There is no support line. When the source host changes something, it is my afternoon. Every tool I add is a tool I now own.
  • Scope discipline is the real job. The test framework can absorb every idea anyone has. I keep a rule: a tool must answer a question someone asks repeatedly, or it does not ship.
  • It grows warts. The server started as one file and is now several thousand lines. I am splitting it into route modules, which is a perfectly normal growing pain and also a tax I would not pay for a bought product.
  • Bought tools have real strengths. Polish, accessibility work, SSO and audit features, someone else carrying the pager. If your question is generic, buy the answer.

My rule of thumb: buy the generic question, build the specific one.

Takeaway

A dashboard is only as good as the question behind it. QA wants to know what is broken and why. Developers want to reproduce it. Leads want the trend and the debt. PMs want to know a ticket is ready, and what was decided in the meeting. One site, one data model, a few hours per idea. That is what ownership buys you.

Next: 6. From laptop tool to team service, on making all of this survive restarts and be safe for a whole team to use.