Skip to main content

The healthy cell that no request could reach

· 23 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the difference between running a system and changing one.

Slatebridge is a fictional cloud dispatch platform for independent home-services contractors — the heating, plumbing, and electrical companies with four to sixty technicians that arrive at your door in a small van. Roughly 2,400 fictional contractor companies and 19,000 technicians use it.

Slatebridge Desk is the browser app where a dispatcher builds tomorrow's route. Slatebridge Field is the phone app where a technician opens that route at shift start and marks each job done on the customer's driveway. Underneath, Slatebridge runs five cells, cell-01 through cell-05, with one shared entry layer in front of all five: Gatehouse.

Beatriz Quintero was the staff platform engineer who owned Gatehouse and Tenant Control, the service that creates contractors and decides where they live. Ivo Krastev was the site reliability engineer on call the morning this story happens. Wes Halloran ran support, so he learned the size of a problem before the dashboards did.

Slatebridge promised contractors one thing above all: from 06:30 local time onward, a technician can open the day's route in one tap. A technician who cannot open a route does not start driving.

How a route normally reaches a technician​

Loading the normal morning journey…

A is not a shard of a database or a slice of a server. It is a complete copy of the product: its own API servers, its own relational database, its own job queue. cell-03 could serve its contractors with the other four switched off.

Each cell owns a disjoint list of contractors, split on contractorId. Gatehouse is what makes five cells look like one product — the industry name for this layer is a cell router. It checks the session, works out which cell owns that contractor, and forwards the request; cell-03 reads the route out of its own database. Two hours later the technician marks a job complete, cell-03 writes that too, and the dispatcher sees it turn green.

Nothing in that path is unusual. It had been running for two years.

Two planes, and who owns what​

Two very different kinds of work live in that picture. The first is the daily business of serving requests from state a component already holds — reading today's route, writing a completion, projecting it onto a board. That is the , and at Slatebridge it lives inside the cells.

The second is the machinery that changes the arrangement: creating a contractor, provisioning a cell, moving a contractor between cells, holding the record of who lives where. That is the , and here it is Tenant Control.

OperationPlaneOwner
Create a contractor, provision a cell, migrate a contractorControl planeTenant Control
Hold the record of which cell owns which contractorControl planeTenant Control
Return today's route from state the cell already holdsData planeThe owning cell
Record a completed job, update the dispatcher boardData planeThe owning cell
Forward a request to the owning cellData plane — the shared oneGatehouse

Notice the last row. Forwarding a route request is day-to-day service traffic, so Gatehouse is data plane too. What makes it different from a cell is that it is the one piece of data plane every cell's requests share — and that, not its label, sets how far a failure there can reach.

The lookup that fixed a real bug​

Six months earlier, Slatebridge had moved a contractor from cell-02 to cell-04. Gatehouse cached each contractor's cell for fifteen minutes, so for several minutes after cutover some instances still held the pre-move answer and wrote job completions to cell-02 — the cell that no longer owned that contractor. The writes were not lost. They were in the wrong database, invisible to the dispatcher board, which now read cell-04. Wes's team repaired eleven fictional job records by hand.

Beatriz fixed it the way you would fix it. Gatehouse stopped trusting a long cache and started reading the answer from the control plane, authoritatively, on the request itself:

GET /v1/placement/c-40711
{
"contractorId": "c-40711",
"cellId": "cell-03",
"placementGeneration": 7,
"asOf": "2026-06-02T09:14:03Z"
}
async function chooseCell(request: FieldRequest): Promise<CellId> {
const placement = await placementApi.get(request.contractorId, {timeoutMs: 250});
return placement.cellId;
}

A 90-second per-instance cache stayed in front of it, which is why nobody thought of this as a control-plane call any more. When the lookup failed, Gatehouse returned 503 PLACEMENT_UNAVAILABLE rather than guessing — also right, because forwarding a write to a guessed cell is exactly the bug Beatriz had eliminated.

Split writes went to zero and stayed there for six months. The fix also meant a Gatehouse instance with a cold cache could not choose a cell at all without a successful call into the control plane.

The morning of the wave​

Slatebridge had just provisioned cell-05, and Tenant Control had a scheduled wave to fill it: migrate 340 contractors out of the four older cells. Each migration copied the data, then held one of the placement store's database connections while it wrote the new cell and verified both sides.

That last detail is the hinge. Connections to that store come from a fixed pool of 120, so while migration work holds them, every other reader — Gatehouse included — waits for one to come free. A read that took four milliseconds now takes however long the queue takes.

It ran at 06:40 because the team's peak dashboard showed traffic peaking at 08:00 — and that dashboard tracked dispatcher sign-ins. The technician peak, 06:45 to 07:15, lived on a dashboard nobody consulted when scheduling maintenance.

Time (ET)Control planeCellsWhat a technician saw
06:40The wave starts, 340 contractorsHealthyNormal
06:44Migration work takes most of the pool; headroom for other readers collapsesHealthyNormal
06:47—HealthyGatehouse scales out; new instances have empty caches
06:50The median placement read passes 250 ms; the 99th percentile reaches 3.4 sHealthyLookups start timing out
06:51—Healthy503 PLACEMENT_UNAVAILABLE; 71% of shift-start requests fail
07:09The wave is pausedHealthyRecovering; normal by 07:13

The 71% is not a tail effect: seven in ten shift-start requests landed on an instance with no cached entry for that contractor, and by 06:51 almost none of those lookups came back inside the budget.

Twenty-two fictional minutes, and the window caught the Eastern-timezone shift start: about 2,700 route loads attempted, 1,900 failed, 260 missed appointment windows across 140 contractors, 300 calls into Wes's queue.

The control plane was never down. It answered every request in that window. It answered slowly, which for a caller with a 250 ms budget is the same thing.

Loading the incident sequence…

Every cell was fine​

This is the part Ivo had to establish first, and the part a reader should refuse to take on trust.

SignalValue during the window
Success rate for requests that reached a cell99.98%
Cell database CPU, all five cells11–14%
Cell job-queue depth0
Per-cell deploy canary, which bypasses Gatehouse entirely1,320 calls, 0 failures

The canary is decisive, and it mattered that Slatebridge had not invented it during the incident. It already existed to verify cell deployments: it calls each cell's internal endpoint with a fixed contractor and a fixed cell, so it never asks the control plane anything. That is why Ivo could rule out five cells in three minutes.

So every route, technician record, and job row sat in a healthy database with spare capacity, and could not be handed to the person driving toward it.

Now the uncomfortable comparison. Cells exist to limit blast radius — the set of customers and resources one failure can affect — and they had been doing it: a bad deploy to cell-02, a corrupt row, a runaway query had each stayed inside one compartment. But every request crossed Gatehouse, and every cold Gatehouse request crossed the control plane. Compartments narrow the blast radius of failures inside a compartment. A cell-based system is only as compartmentalised as the least compartmentalised thing on its serving path.

The lever was inside the failure​

At 07:02 Ivo opened the runbook. Its first mitigation was sensible on paper: re-point the affected contractors to a spare cell. Re-pointing a contractor is POST /v1/placement/{contractorId} — a write to the same store that was already saturated, queued behind 340 migrations. The recovery lever was inside the thing that had failed.

What worked was subtraction. The migration workflow ran from its own separate store, so Ivo could pause it without touching placement, and did at 07:09. Lookups were healthy four minutes later.

Cheaper ideas came up in the channel. Being precise about which ones work is most of the lesson.

IdeaWhat it does
A bigger store or connection poolMoves the threshold, leaves the dependency
Retry the lookupAdds load to a pool that is already full
Raise the 250 ms budgetTurns a fast failure into a slow one, and fills Gatehouse's own request slots
Move the wave to midnightSaves this morning, changes nothing — the next cold-cache instance during any control-plane slowdown finds the dependency again
Give serving-path reads their own pool or replicaWorks, and cheaply. Migration writes can no longer starve them, so this incident does not recur — but the ownership answer still comes from outside the cell. A slow Tenant Control deploy brings the same morning back, and a lagging replica is worse than that: it answers, promptly, with a placement that has been superseded, which is the split-write bug the lookup existed to prevent
Forward to any cell when placement is unknownPuts completions in a database that does not own them

That last row has a much better-looking cousin: keep serving the 90-second cache entry you already have. During a wave touching 340 of 2,400 contractors it is right for about six in seven, which is exactly why it appeals. It is also correction A minus its most important part, and it swaps a loud availability bug for a quiet correctness bug that surfaces as a completion nobody can find. Serving a stale answer is safe only once something can authoritatively refuse work it does not own. Nothing could, yet. That ordering is the correction.

The old path was:

request arrives
-> ask the control plane which cell owns this contractor
-> forward
-> control plane unreachable -> fail the request

The corrected path needed a different first step and a new last one:

request arrives
-> read the cell from a map the control plane filled in advance
-> forward
-> cell confirms it still owns the contractor, or names the one that does
-> control plane unreachable -> keep serving the last published placement

Correction A: publish the map before it is needed​

Beatriz's default correction was to stop asking.

Every 30 seconds, Tenant Control writes a complete placement snapshot to object storage: every contractor, every cell, a version, a timestamp. Each Gatehouse instance loads the newest one before accepting traffic and refreshes in the background. The serving path reads memory:

function chooseCell(request: FieldRequest): CellId {
return placementMap.get(request.contractorId) ?? hashToCell(request.contractorId);
}

The hash fallback covers a contractor too new to appear in any snapshot: it lands on some cell, and that cell says who really owns it. One extra hop, never a wrong write.

An instance that cannot refresh does not fail requests. It raises a snapshot-age number and keeps serving the last answer it has. That property has a name: . The system keeps working through a dependency's impairment because nothing has to change for it to keep working.

One dependency survives, and the story owes you it: a booting instance still needs a snapshot, so it cannot enter service without reaching object storage. Scale-out is therefore still exposed — which is what triggered this incident. The dependency moved rather than vanished, from a store migrations were saturating to one that changes every 30 seconds and serves nothing else. A team that wants scale-out to survive even that bakes the last snapshot into the image.

The snapshot is 180 KB: cheap to hold, expensive to trust. The map can now be wrong, so it cannot be the authority any more — and that part is required, not optional. A stale map with nobody checking it trades an availability bug for a correctness bug.

So each cell became the authority for its own contractors. Each keeps an ownership record carrying a generation — a counter that only goes up, so the higher number wins.

Cells load that same published snapshot too, which is what lets a cell that does not own a contractor still name the one that does: 409 WRONG_CELL. Gatehouse follows that redirect once and updates its own entry. If the second cell disagrees too, it fails the request rather than bouncing it between them.

The map became a hint. The cell became the authority. That sentence, not the snapshot, is the architectural change.

Migration cutover changed shape too:

  1. Start replicating the contractor's data to the target cell.
  2. Put the source cell's contractor into read-only handoff. Completions get 409 HANDOFF_IN_PROGRESS and a retry-after; route loads are still served.
  3. Wait for the drain — in-flight writes finished, and the target caught up to the source's final state.
  4. Write the authoritative record with an incremented generation, and publish a new snapshot.
  5. Activate the target at the new generation, and wait for it to acknowledge.
  6. Release the source, once that acknowledgement is in and the snapshot naming the target has actually been published.

Step 3 is not padding. The copy starts while the source is still taking writes, so if the generation flips before replication catches up, a completion written in between ends up in a database nobody reads — which is the bug this whole design exists to prevent, rebuilt by the design meant to prevent it.

Step 6 has two conditions rather than a stopwatch, and both were earned the hard way. Waiting for every Gatehouse instance to confirm sounds safer and is worse: one wedged instance blocks every migration indefinitely. But a bare timer is not safe either — it can fire while the snapshot naming the target is still unpublished, and then the released source's own copy still names itself, it has no new owner to point at, and a healthy contractor becomes unreachable. So release waits on the two things the migration cannot finish without anyway: the target saying it is live, and the new placement actually being published. Neither of those is a router.

That is what makes the redirect dependable — not that a released cell is fresh, but that a cell is not released until the answer it will redirect to exists. And this happens routinely, not rarely: the wait is about as long as one refresh cycle, so an instance that misses a single fetch is already stale when release fires. Meeting a released cell is the ordinary case. It costs one extra hop, and the second exercise below is built to provoke exactly that.

Now the bill. Cutover is no longer instant: one contractor goes read-only, and a technician feels it as a completion rejected for about ninety seconds. That ninety seconds is the healthy case — one publication to name the target, one more to release on — and it assumes steps 2 to 6 all complete. Lose Tenant Control between steps 2 and 4 and that contractor stays read-only for the length of the impairment. A stalled snapshot writer does the same thing, and deliberately: release will not fire until the new placement is published, so the migration stalls instead of finishing into a fleet that cannot route to it. Correction A moves the risk from every contractor for twenty-two minutes to one contractor for the length of an outage: better, and not free. Five cells also each carry ownership logic that did not exist before, and somebody operates a snapshot writer whose failure mode is silent staleness — which is why snapshot age became a number the team commits to.

Correction B: bind the cell into the session​

Slatebridge's alternative went the other way: if the shared layer should not have to ask, do not make it the one who knows.

Sign-in is served by the owning cell, and it mints a signed session carrying the cell and the generation. Slatebridge Field then addresses that cell's own endpoint directly. In the steady state nothing shared is consulted at all, and the cell alone decides whether it owns the contractor.

The session lifetime is the design, so it has to be named: 30 days. That is what makes the steady state real — a technician signs in once a month, not once a shift, so the bootstrap path carries a trickle instead of the whole 06:45 peak. Choose hours instead and every technician bootstraps every morning, the bootstrap map is back on the critical path for exactly the traffic that failed, and correction B collapses into correction A with five extra endpoints.

Thirty days is also the bill. You cannot re-point a contractor on demand: draining a cell means waiting up to a month or forcing 19,000 technicians to sign in again mid-shift, and a leaked session stays bound to its cell until it expires, so revocation becomes per-cell work. Five cells need five public endpoints, five certificate rotations, and five sets of rate limits — the duplicated per-cell operation you accepted when you chose compartments, now visible at the edge. And a first sign-in has no session and no cell, so a small statically stable bootstrap map survives: correction B shrinks correction A rather than deleting it.

Correction A is the default. Correction B is what you choose when you would rather pay in operations than in staleness.

Concepts in this story4 concepts
Control planePlatforms and control planes · Plane boundaries

The machinery that changes a system: creating, describing, updating, deleting, and listing resources, and propagating those changes to wherever they take effect. Placement, provisioning, migration, and credential issuance are control-plane work.

Data planePlatforms and control planes · Plane boundaries

The daily business of serving requests from state a component already holds. A data plane is usually simpler than the control plane that arranges it, which is why it can stay healthy while the control plane is degraded.

CellPlatforms and control planes · Fault isolation

A complete, independent instance of a workload that shares no state with other cells and serves a subset of customers or resources. Cells narrow the scope of failures that happen inside a cell; they do not isolate a dependency every cell’s traffic must cross.

Static stabilityPlatforms and control planes · Availability under impairment

The property that a system keeps working correctly while a dependency is impaired, because nothing has to change for it to keep working. It is bought with pre-provisioned capacity or pre-published state, and paid for in staleness, cost, or operational duplication.

What the game days proved​

Correction A has two risks, and one exercise cannot test both. Making the control plane unavailable proves serving survives — but while it is unavailable nothing can change ownership, so a stale map is trivially correct and the staleness risk goes untouched. Slatebridge ran two exercises. Both sets of numbers are clearly fictional.

Exercise 1 — the placement store made unavailable for thirty minutes on a weekday morning.

MeasureFictional result
Route loads served100%
Worst placement-snapshot age31 minutes
Migrations completed0

The last row is the honest one. Static stability kept serving alive. It did not keep administration alive, and was never supposed to: for thirty minutes nobody could onboard a contractor, provision a cell, or finish a migration, because those are changes.

Exercise 2 — Gatehouse's snapshot refresh stalled at five minutes while a 60-contractor wave ran. The control plane stayed healthy and kept publishing every 30 seconds, so only the routing layer fell behind. Sources were released exactly as step 6 says — on the target's acknowledgement and a published snapshot, never on confirmation from the routing layer.

MeasureFictional result
Route loads served100%
Requests sent to a cell that no longer owned the contractor12
Of those, requests that wrote to the wrong cell0
Worst read-only window for a migrating contractor71 seconds

This is the one that earns the design, and it is worth walking. A stalled Gatehouse kept routing twelve requests to cells that had already handed those contractors over. Each of those cells had a current snapshot, because only the router was stale — so each could name the new owner, Gatehouse followed the redirect once, and the request was served. Nothing was written to a cell that did not own it. That is the cell-as-authority claim tested rather than asserted: the map was wrong, and being wrong cost one extra hop.

Two things neither exercise covered, and both are the story's own admissions rather than oversights. Neither made the snapshot store unreachable while Gatehouse was scaling out — the exposure named back in correction A, and the closest analogue of the morning that started this. And neither covered a control-plane failure during a cutover, the stranded-contractor case. Both are still on the list.

Coming out of a degraded window needs no coordination, which is the quiet payoff of moving the authority: instances resume at whatever snapshot they can reach, and any disagreement resolves the way it does on an ordinary morning — the owning cell decides.

One measurement changed too. Failures had been counted in a single number, which is why the first five minutes were confusing: the graph said the platform was failing and could not say which plane was. They are now counted as failed before reaching a cell and failed inside a cell, and the top-level health check no longer runs through Gatehouse alone.

Here is where responsibility ended up, in both corrections:

Loading the corrected production topology…

The two designs disagree about who holds the answer: a shared layer reading a published map, or a client holding a signed one. They agree about what matters. The control plane is off the serving path, the cell is the authority on what it owns, and healthy capacity keeps serving when the machinery that changes it stops.

Transfer the question​

Nothing here is about dispatch software, and nothing requires cells.

Picture a payments platform that resolves each merchant's active plan before processing a charge. The plan is administration data — sales changed it, support can change it, it lives in the admin database — and now it is on the charge path. When the plan service slows, every healthy processor stops charging healthy merchants whose plans have not changed in a year. Or a build system whose runners ask a central scheduler where the shared cache lives: warm runners, healthy cache, no build starts.

Three questions find these before they find you.

  1. Which questions does my serving path ask that only a control plane can answer?
  2. If that answer never came back, would healthy capacity keep serving — and how stale would the last answer be allowed to get?
  3. Who owns "this work is mine" closely enough to the work that the answer survives the control plane?

And one for the worst hour rather than the ordinary one: is my first mitigation step itself a control-plane operation? If your runbook opens with a call to your own administrative API, rehearse it while that API is slow, not while it is healthy.

The rule Beatriz wrote at the top of the runbook afterwards was short:

Running a system and changing a system are different jobs. Anything that has to change before a healthy component can serve is on the serving path, whichever side of the diagram you drew it on.

Evidence and fiction note​

The vocabulary here is not invented. AWS's Fault Isolation Boundaries whitepaper (2022-11-16), in its section on control planes and data planes, defines a control plane as the administrative APIs that create, read or describe, update, delete, and list resources; defines the data plane as what provides the primary function of the service; and says data planes are intentionally less complicated, which makes failure statistically less likely there. AWS Well-Architected's REL11-BP04 adds that data planes handle day-to-day service traffic — which is why this story calls a shared entry layer data-plane work — and carries the recovery point: use a minimal number of control-plane operations when recovering, and treat reliance on extensive control-plane actions for remediation as an anti-pattern. Static stability using Availability Zones, by Becky Weiss and Mike Furr, supplies the definition used here: in a statically stable design the overall system keeps working even when a dependency becomes impaired. Its own examples are about pre-provisioned compute capacity, not placement maps, and it never claims the property is free.

The cell material is AWS Well-Architected guidance, Reducing the Scope of Impact with Cell-Based Architecture (2023-09-20). Its page on what a cell-based architecture is defines a cell as an isolated instance of a workload sharing no state with other cells, assigns provisioning and customer migration to the control plane, and calls the routing layer the thinnest possible layer. Its page on the resilience of the cell router says plainly that the router is the only component holding the shared state of all cells and presents itself as a single point of failure. And it already describes correction A as a real design: in one documented router option, the control plane writes the cell mapping to object storage and the mapping lives in memory on the router.

None of those sources describes an incident at any company, and none supports any figure above. Everything else is fiction: Slatebridge and its two apps, Gatehouse, Tenant Control, Beatriz Quintero, Ivo Krastev, Wes Halloran, every cell name and identifier, the 30-day session lifetime, and every count, latency, rate, duration, and exercise result. No named database, queue, object store, or cloud service is claimed to behave in any particular way; the shared placement store is described only as a single regional relational database with a fixed connection pool, which is ordinary engineering rather than a vendor claim. This is an architecture lesson, not operational advice for any specific platform.