The Empty 200
A rejected request that returns a bare HTTP 200 looks, in a log, exactly like a served one. That is worse than an error page, because nothing flags it for investigation. Notes on the bug class I now check for first.
I built the multi-tenant plumbing this week, before having any tenants to put on it. The alternative is migrating a live workflow out from under a paying customer, and I would rather spend a few days on plumbing than have that week.
The plumbing produced a bug I have thought about more than the plumbing itself: a rejected request that returns HTTP 200 with an empty body.
The setup
Every client-facing workflow first calls a small sub-workflow that takes a tenant ID and answers one question: who is this, and are they allowed to be here? It returns either the tenant’s configuration or a structured refusal — unknown tenant, paused tenant, missing ID, bad config. None of those refusals is thrown as an exception. A paused client is an expected outcome, and treating it as a crash would take down every other client’s traffic with it.
The sub-workflow is three nodes: look up the row, decide, return.
The bug
The workflow engine skips any node whose input has zero items. When the lookup matched no rows — an unknown tenant, the case the resolver exists to catch — the node holding the fail-closed logic never ran, nothing reached the output, and the caller got back a bare HTTP 200 with an empty body.
In a log, a rejected tenant returning 200 OK is indistinguishable from a served one. There is no error to alert on, no non-2xx to graph, no exception to catch. The system reports perfect health while serving nobody. An error page would have been better, because an error page gets investigated.
The fix was one setting: force a placeholder item through on a zero-match query, then filter the placeholder out before deciding whether a real row was found. I re-ran the unknown-tenant case and got a 404 with a structured body.
Why I am writing it down
A week later I hit the same bug in a different workflow. The error handler’s retry-budget lookup had no row on a workflow’s first failure, so the node was skipped and the failure produced no retry and no alert. The first time something broke, the component that tells me something broke was itself broken.
Both cases have the same shape: a lookup that returns nothing, followed by a node that decides based on what the lookup returned. Both passed static validation, because nothing was malformed. The bug exists only at runtime, only on the empty branch, and the empty branch is the one nobody tests.
$ curl -s -o /dev/null -w '%{http_code}' $ENDPOINT -d '{"tenant_id":"nope"}'
That two-second call is the whole test. I now run some version of it against every lookup-then-branch I build, and “what does this return when it finds nothing” is the first question I ask rather than the last.
The general form: any system where the negative path is quieter than the positive path will eventually report health it does not have. It errs toward reassurance, which is the direction you are least likely to check.
The full write-up, including the registry design and the four rejection cases, is on the Marain site.