Series: 1. Beyond the test suite · 2. Rules first, LLM second · 3. Comparing responses across a flag · 4. Making the test framework agent-native · 5. Dashboards for the whole team · 6. From laptop tool to team service
Every internal tool has a moment where it stops being yours.
For a long time, my test framework was a thing I started in a terminal, used, and closed. If it broke, I knew why. If its data vanished, I could re-pull the runs. A tool like that is a notebook, not a service.
Then somebody asks for the link, which is a lovely request and a terrifying one: your laptop is now part of someone else’s workflow.
This article covers what it takes to turn a local dev tool into a hosted, restart-safe service for a team: the design choices that paid off, four gotchas worth knowing, and the guardrails. I keep the hosting details deliberately high level. The lessons transfer; the specifics stay out of it.
Two modes, one codebase
In development the test framework is two processes. The front end runs on one port with hot reload and proxies API calls to the back end on another. That is the usual React-plus-Express arrangement and it is pleasant to work in.
Hosted, it is one process. The server serves the built front end and the API from a single port. No proxy, no second container, no CORS puzzle. One thing to deploy, one thing to health-check.
That split sounds trivial, and it is, as long as you decide early that the two modes must be the same code. The moment hosted mode needs a special build of the server, you are maintaining two products.
The storage adapter: the change that made everything else possible
On a laptop, state lives on disk: archived reports, cached analysis, comparison projects, load-test notes. Files in folders. Easy to inspect, easy to delete.
A container on a cluster is the opposite. It can be killed, rescheduled, or replaced at any time, and whatever was on its disk goes with it. Restart-safe means the state lives somewhere else.
The tempting move is to sprinkle “if hosted, talk to the cloud object store, else use the file system” through every route. I have seen that movie. It ends with forty if statements and a bug that only appears in production.
Instead I put one small interface between the routes and the bytes:
Storage
getJson(key) -> object | null
putJson(key, object)
putObject(key, bytes)
list(prefix) -> keys
deletePrefix(prefix)
copy(from, to)
Two backends implement it:
- Local disk, the default. This is what development uses, and it is what the tests use.
- Cloud object store, selected by configuration when hosted.
flowchart TB subgraph devMode ["Development"] FE["Front end, hot reload"] -->|proxy| API["API process"] end subgraph hostedMode ["Hosted"] ONE["One process serves the front end and the API"] end API --> STORE["Storage interface"] ONE --> STORE STORE --> DISK["Local disk"] STORE --> CLOUD["Cloud object store"]
The routes never learned which one they were talking to. When the test framework became a stateless container, the routes did not change. The work was writing a second implementation of six methods, not rewriting the application.
A side benefit I did not plan for: the interface is small enough that “what does this feature store, and where?” has a one-screen answer. Nothing audits your data model like being forced to describe it as six verbs.
Container choices that were not accidents
The image looks boring on purpose. Three decisions deserve a sentence each.
A glibc base, not Alpine. The server shells out to Playwright, and Playwright drives real browsers. Those browsers need a standard glibc userland. The slim musl-based images that are lovely for plain Node services are the wrong tool here. I use the vendor’s Playwright image as the base and accept the weight. The finished image is around 0.9 GB. That is not elegant. It is also pulled once per node, not once per request.
Base tag pinned in lockstep with the library version. Playwright the npm library and Playwright the browser bundle must agree. If the library says one version and the image ships browsers for another, you get failures that look like test problems and are actually plumbing. So the image tag and the library version move together, in the same change, always. A mismatch should be impossible to merge, not merely discouraged.
Runtime layout mirrors the repo. The server finds the test tree relative to its own location. In development that is just where the folders happen to be. In the container I could have “tidied” the layout and then spent a day wondering why the runner found no tests. So the image recreates the repo’s directory shape. Boring, predictable, correct.
Two smaller notes. The process runs as a non-root user. And dev dependencies stay in the image, because the server runs TypeScript directly through a runner with no compile step. If you ship without the runner, you ship without a server. I would rather keep the dependency and delete a build stage than the other way around.
Four gotchas worth knowing
None of these is exotic. That is the point. They are the ordinary, unglamorous things that bite when software moves house.
1. Environment loading order with ES imports
The natural place to load environment variables from a file is the first line of the main module. With ES modules that is too late: all imports are evaluated before the importing module’s own body runs, so any imported module that reads a variable at import time sees undefined.
The fix is to put environment loading in its own tiny module and import it first:
// load-env.js (imported before anything else)
import 'dotenv/config';
// server.js
import './load-env.js'; // evaluated first, in order
import { startServer } from './app.js';Imports are evaluated in order, so the env module finishes before the next import starts. If your config “randomly” works locally and fails elsewhere, check this before anything else.
2. The explicit-credentials pitfall
Cloud SDKs have a built-in credential chain that tries several sources in order, which is what lets a hosted workload run without keys in its config.
Passing an explicit credentials object when you construct the client short-circuits that chain. It can work on a machine that happens to have keys in its environment and then fail in the place where the platform was supposed to supply them.
Lesson: when a library has a smart default for something sensitive, do not override it “to be explicit”. Explicit is not always safer. Sometimes it simply switches off the mechanism that would have worked.
3. The hostname collision
Container platforms typically set an environment variable named HOSTNAME to the name of the running instance. If your tests read HOSTNAME as the address of the API under test, as mine did, the collision is invisible locally, where it is unset or harmless and the suite falls back to its default. In the container it holds a generated instance name, every request goes to a host that does not exist, and you get a whole run of connection errors from a variable that seems unrelated to tests.
What to watch for: generic names like HOSTNAME, HOST, PORT and USER are shared territory, and the platform may set them. Give your own settings prefixed, project-specific names that cannot collide.
4. “0 tests in 0 files” is a green-looking nothing
This one is my favourite, because it is the same disease as the empty type-check I described in article 2.
Some tests read a secret when the file is imported, not when the test runs. If that secret is missing, the file throws during collection. The runner does not fail loudly. It reports that it found 0 tests in 0 files and exits quietly, and everything downstream that watches the exit code says: fine.
Nothing ran, and nothing complained. A run that examined nothing is indistinguishable from a run that passed, unless you check how many things it examined.
The defence is to treat an empty result from the runner as a failure of the instrument, not a success of the code: compare the count you expected with the count you saw, and never call zero a pass. Check the denominator. A green you cannot count is not a green.
Guardrails: the point where “my tool” becomes “our tool”
A tool on your laptop spends your quota and touches your data. A tool on the cluster spends everyone’s. Once there is a team, the question changes from “does it work?” to “what is the worst thing it can do on a bad day?” The guardrails are my answers.
A per-user daily AI quota. The AI features cost real money per call, and a curious user in a loop could burn a lot of it. Each user has a daily request ceiling, and the remaining budget is visible in the health endpoint and response headers, so nobody is surprised by a wall. The counter lives in the process and resets on restart. I accepted that deliberately: the real cap is the spend limit at the LLM provider, and the in-process counter is there to stop honest mistakes, not to be an accounting system. Know which control is the backstop and which is the speed bump, and do not build the speed bump to backstop standards.
A hard floor on calls to external sources. Anything that polls an outside system, such as the dependency-alert refresh from article 5, has a minimum spacing between live calls, no matter how many people press refresh. Ten impatient users should produce one upstream request, and the rest get the last good answer. Backoff on rate-limit responses and serving cache on failure live in the same family. The principle: the number of calls you make to someone else’s system must not scale with the number of people who are looking at your page.
A QA-only target, by design. The test framework points at the test environment. It is a testing tool and it behaves like one. I have planned startup protections that refuse to boot at all if it is ever configured against production credentials. I would rather a misconfiguration be a crash at start than an incident later.
Secrets injected at runtime, never baked into the image. The image contains code and nothing you would mind leaking. Configuration arrives when the container starts.
A stuck-run watchdog. Test runs are long-lived child processes. If one hangs, it should not hold a slot forever, so a watchdog notices a stuck run and cancels it.
None of these is clever. All of them came from asking the same bad-day question about a different feature.
The rollout: boring on purpose
I did not move the test framework in one heroic weekend. I staged it, and each stage was verified before the next began:
- The image. Get the container to build and run the tests on its own. No hosting yet.
- The server refactor. Make the server serve the front end and API as one process, and untangle what assumed a laptop.
- Storage. Introduce the adapter, add the cloud backend, prove state survives a restart.
- The AI limit. Quota and headers in place before real users could reach the AI features.
- Deploy wiring. The container actually running on the cluster, with configuration injected.
- The access gate. Only then, open the door to the team.
The order is the lesson. Notice that the AI limit lands before the door opens, and the gate lands last. Anything that protects against misuse should exist before the users do. Each stage had a plain question it needed to answer (“does state survive a restart?”, “does an empty run fail?”) and I did not move on until I had seen the answer, not assumed it.
It is slower than flipping a switch. It aims for a dull first day of real use, and dull is the highest compliment a rollout can get.
What I would tell you to steal
- Put one small interface between your routes and your storage, and make the default backend the dumbest one.
- Pin things that must agree, and pin them together.
- Do not use generic environment variable names for your own meanings.
- Let the platform’s credential chain work. Do not override it for the sake of being explicit.
- Make an empty result a failure. Count what you examined.
- Build the guardrails before the users arrive.
Looking back at the whole series
Six articles ago I made a claim: writing tests is no longer the scarce thing. AI can produce test cases faster than anyone can read them. What remains scarce is understanding what the results mean: why hundreds of tests failed, what changed when a flag flipped, which failures are the same failure, who needs to hear about it.
Everything in this series is an answer to that. The triage pipeline uses deterministic rules because most failures are boring and a rule is cheaper, faster, and more trustworthy than a model, and uses the model only for what is left, with every bit of its bookkeeping checked. The response comparison turns “does the migration behave identically?” from a feeling into a table. The MCP layer means those same capabilities are available to an agent in conversation, without duplicating a line of logic. The dashboards let QA, developers, leads, and product people each ask their own questions of the same data. And this last piece is the unglamorous part that makes any of it usable by someone other than me.
A few threads ran through all of them:
- Deterministic first, judgement second. Reserve the model for the part that actually needs judgement.
- State what you examined. A green with no denominator is the most dangerous result a tool can give you. It can show up in a type-check, in a test runner, and in a model’s bookkeeping, and it will show up in yours.
- Humans approve mutations. Propose, approve, apply, verify.
- Build the tool the third time you ask the AI for the same report. Cheaper, steadier, and the AI can call it afterwards.
And a quieter one: owning the test framework changed the questions I thought I was allowed to ask. Building gives you your own questions, in hours, over your own data model. The price is maintenance, and you should be honest about that price. Keep the scope tight, split the files when they get fat, and delete the features nobody opens.
If you take one thing from six articles, take this: the interesting work in test automation has moved downstream of the test run. Go build the thing that tells you what the run meant. Then, when somebody finally asks for the link, make sure it can survive being theirs too.
This is the last article in the series. Start again from the top: 1. Beyond the test suite.