Skip to main content

The green test that had never been told the rule

· 20 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the difference between a test that records output and a test that states a rule.

Kestrelway Transit is a fictional regional transit operator. Riders tap a bank card or a phone at a gate; there is no ticket to buy and no fare to choose in advance.

Kestrelway built FareBridge to price one rider's complete service day, which runs from 03:00 to 03:00. It takes everything a rider did in that window and produces one figure.

Ola Fadeyi was the fare policy analyst, and owned the published fare rules, their wording, and the date each took effect. Hugo Bertrand was the fare platform engineer who owned FareBridge, its regression suite, and the gate that decided whether a change could ship. Grace Mbeki led revenue operations and owned daily settlement reconciliation.

FareBridge decided what a day cost. Settlement decided that the charge was sent — once.

How a tap becomes one charge​

Loading the normal fare journey…

No rider ever typed a fare. Two facts in that path carried a rider category, and they were not the same fact. The card product was a historical observation: the fare product registered against that card token, sometimes years earlier. The verified entitlement category, resolved by a separate service at the end of the service day, was the authoritative answer to a question about today.

How those two can disagree matters later, so it is worth stating now. A concession — Kestrelway's reduced-fare eligibility, granted for age, disability, or an income scheme — is granted by registering a concession card product, and the entitlement service then re-verifies it periodically and can let it lapse. So a card product can go stale in exactly one direction. It can still say CONCESSION after the entitlement behind it has expired. It can never say STANDARD for a rider entitled to a concession, because becoming entitled is what registers the card.

The fare product was deliberately plain. Amounts are in minor units — the smallest unit of the operator's currency, as cents are to a dollar:

ElementAmount (fictional)
Base fare per journey250 minor units
Peak surcharge per journey starting inside a peak window75 minor units
Daily cap, STANDARD700 minor units
Daily cap, CONCESSION350 minor units

One detail matters. The concession benefit was delivered only through the lower cap. Base fares and surcharges were identical for every rider, so the only thing a rider's category could change was which cap applied.

The golden days​

In 2024 Hugo's team captured sixty complete rider-days from production — messy ones, with reversed taps, gate errors, midnight crossings, and combinations nobody would have written by hand. For each one they recorded the figure settlement actually used, totalChargedMinor, and approved it. The team called them the golden days, and trusted them more than anything else in the repository.

That was a good decision. Over two years the golden days caught a journey-pairing regression, a service-day boundary bug, and a botched cache change, at almost no maintenance cost, over input shapes no hand-written example would reach. Nothing had disturbed them in between: CAP-4's capping expression was already in force when the corpus was recorded, and the rule set published in April 2025 carried it forward untouched.

Recording current output and comparing future runs against it is a ; the tooling world calls the same idea snapshot, approval, or golden-master testing. This story calls the corpus the golden days and one entry in it a recorded expectation.

The gate was three lines:

test.each(riderDayIds)('rider day %s', (id) => {
expect(priceServiceDay(riderDay(id)).totalChargedMinor).toMatchSnapshot();
});

The mechanics are what the documentation describes. The first run records the expectation; a later mismatch fails; --updateSnapshot in Jest, --update in Vitest, or -u in either, rewrites the approved value; the artifact is committed and reviewed with the code. Neither runner writes snapshots in a continuous-integration run by default, so nothing was being quietly regenerated in the pipeline.

Jest's own documentation is candid about the risk in that workflow: the goal is to "fight against the habit of regenerating snapshots when test suites fail instead of examining the root causes of their failure."

Hugo's team did review them.

What the green suite proved​

ClaimEstablished by a green suite?
Each recorded rider-day's total matched an earlier observationYes
Journey pairing, service-day boundaries, and surcharge counting were unchanged for those sixty inputsYes
No refactor had silently altered pricing for those sixty inputsYes
The totals were the totals Kestrelway's published rules called forNo

The first three were real production guarantees. None was fake, stale, or bypassed; the suite did exactly the job the tooling documents, which is to make sure output does not change unexpectedly.

The fourth row had never been in scope. Nobody decided that. It had simply never been written down anywhere.

The rule that changed​

On 24 June Ola published fare-rules.2026-07-01, effective at 03:00 on 1 July. It changed rule CAP-4:

What CAP-4 saysv1v2
Charge for the daymin(sum(base), cap) + sum(surcharge)min(sum(base) + sum(surcharge), cap)
Peak surchargesSit outside the capCount toward the cap
Cap category comes fromVerified entitlement at end of service dayVerified entitlement at end of service day, stated explicitly for the first time

Two clauses, one publication. The first was the change everybody discussed: surcharges moved inside the cap, so capped riders would pay less. The second was a clarification Ola had written down after a colleague asked about lapsed concessions.

One shared helper​

Hugo implemented v2 on 29 June. In v1 the cap applied to base fares and the surcharge was added on top, so the change collapsed two steps into one.

The new helper was written to be shared with a weekly-cap prototype, and it took the shape that prototype could supply:

applyCap(basisMinor, capContextFromTaps(day.taps));

A tap carries a card token, a station, a time, and the fare product registered against the card. It does not carry an entitlement. Inside the shared helper there was no entitlement to read — and passing one in would have given the helper a second input shape, which is the fork the refactor existed to remove.

Twenty-three of the sixty recorded expectations changed. Every changed total moved down, which is precisely what v2 predicted. Hugo read the diff, accepted it with -u, and committed the rewritten expectations. A colleague reviewed both the code and the expectations. The suite was green. The change shipped.

Two diffs that meant different things​

RD-0217 was a standard rider with three journeys, two in peak:

base 3 x 250 = 750 surcharge 2 x 75 = 150

v1: min(750, 700) + 150 = 850
v2: min(750 + 150, 700) = 700

RD-0431 had four journeys, two in peak. Its recorded input carried a card product of CONCESSION_2023 — and also carried the authoritative field, which said STANDARD. This rider's concession had lapsed before the day was even recorded.

base 4 x 250 = 1000 surcharge 2 x 75 = 150

v1: min(1000, 700) + 150 = 850 the same recorded total as RD-0217

v2 with the verified entitlement category: min(1150, 700) = 700
v2 with the card-product category: min(1150, 350) = 350

The cap category was now derived from the taps, so the verified entitlement was not ignored — after the refactor it was simply not in scope where the category was chosen.

This is the before-and-after Hugo reviewed:

before -> after, as the recorded diff showed it

RD-0217: - 850 + 700
RD-0431: - 850 + 350

Both fell, from the same number. One fell because surcharges now counted toward the cap. One fell because the cap had been chosen from a card instead of an entitlement. The recorded expectation held a total and nothing else: no rule identifier, no cap amount, no category, no rule-set version.

That is a failure of . Every expected value in the golden days came from the system's own output. A recorded observation can tell you that something moved; it has nowhere to record why, so it cannot separate a number that changed for the intended rule from one that changed for an unintended rule.

Notice what the expectation did contain. RD-0431's input said CONCESSION_2023, and 350 is the concession cap. Inside the four corners of the entry under review, the new value was consistent with everything visible. Accepting a recorded diff is a plausibility judgment, not a re-derivation — and re-deriving sixty rider-days by hand is the work the golden days existed to avoid.

Loading the incident sequence…

Money that had already moved​

At 03:00 on 1 July, FareBridge began evaluating v2. Settlement did what it always did: at the end of each service day it submitted one aggregated charge per card. Nobody was inspecting a queue, and nothing yet suggested holding a batch.

On 8 July, Grace's reconciliation flagged a shortfall among the riders who reach the cap. Over eleven service days, a fictional 2,610 rider-days had been priced with the 350 cap when 700 applied, under-collecting a fictional 831,400 minor units (8,314.00). Kestrelway decided not to back-bill riders for an error it had made, and absorbed the shortfall. That number does not come back; no corrected design recovers it.

Reconciliation is worth naming accurately, too: it spoke on the seventh service day of charges. Detection after an irreversible effect is not a control.

Then the correct fix failed the gate​

On 9 July Hugo found the substitution and restored the verified entitlement as the cap category.

Four tests failed — the four recorded expectations accepted on 29 June. The gate reported the correct fix as a regression, because the corpus now asserted the defect.

There was a way through that in minutes, and it deserves naming: revert the 29 June commit. A revert restores the code and its recorded expectations together, so the corpus goes green at once and the gate is no obstacle. Its cost is why nobody took it — reverting meant pricing under CAP-4 v1 after 1 July, charging capped riders more than the published rule now allowed, and an operator would rather under-collect than charge above its own tariff.

So the team fixed forward, and what the two remaining service days bought was a written justification strong enough to overrule the most trusted artifact in the repository. The fix shipped on 11 July.

One mechanism, seen twice. Deleting the golden days would not have stopped the substitution shipping: it compiled, it was reviewed, it was plausible. What the acceptance did was rewrite the only written description of intended capping behavior so that it agreed with the code. After 29 June the repository held a reviewed, committed artifact positively asserting that RD-0431 should cost 350. That is what made the second failure inevitable rather than unlucky, and what separates this from "no test covered it".

The invariant that was true the whole time​

Most engineers, told this story, propose the same fix immediately: assert the rule.

expect(priced.totalChargedMinor)
.toBeLessThanOrEqual(CAP_MINOR[day.entitlement.categoryAtServiceDayEnd]);

Run that against the shipped code and it passes. 350 ≤ 700. Substitute any of the 2,610 affected rider-days and it still passes.

Not because upper bounds are generally blind to a wrong category — a wrong category can pick a cap that is too high or too low, and an upper bound catches one of those. It could only ever be too low here, because of the one-directional staleness established earlier: the only error Kestrelway can produce is generous, and an upper bound cannot see generosity. In an operator where staleness could run the other way — a rider charged against STANDARD when CONCESSION applied — the same assertion fails, and the change never ships.

Which is the transferable question, sharper than "write the invariant": which direction is my invariant blind in?

The assertion that catches this defect in either direction constrains the reason, not the amount:

expect(priced.capCategory).toBe(day.entitlement.categoryAtServiceDayEnd);

That line makes FareBridge say which category it used, and makes the test name where that category must come from. Neither the recorded total nor the upper bound asks either question.

The old shortcut was:

the recorded day totals match
-> pricing is correct
-> release

The corrected path needs the reason in it:

the recorded day totals match
-> every stated rule assertion holds, including which record the cap came from
-> each changed total is attributed to a declared rule change
-> release, with the recorded totals kept as a change detector only

Correction A: put the rule where its owner can read it​

Hugo's default correction moved intent out of the tests entirely. Ola's published rules became a versioned, checked-in artifact with a validity window — an set that FareBridge selected by service date:

{
"ruleSetVersion": "fare-rules.2026-07-01",
"effectiveFrom": {"date": "2026-07-01", "localTime": "03:00", "zone": "operator-network"},
"supersedes": "fare-rules.2025-04-01",
"rules": [{
"id": "CAP-4",
"version": 2,
"capBasis": ["BASE_FARE", "PEAK_SURCHARGE"],
"categorySource": "entitlement.categoryAtServiceDayEnd",
"capMinor": {"STANDARD": 700, "CONCESSION": 350}
}]
}

Two fields do the work. categorySource is the field whose absence caused the incident, now part of the specification and in a file Ola can read. effectiveFrom carries a local wall time with its zone rather than a fixed offset, because a 03:00 boundary is a local-time rule, not an instant: an offset is an hour wrong on one side of a clock change, and old rule versions are exactly the ones re-priced across one.

The suite then drove example rows Ola authored from the rule text, including a lapsed-concession row. Each names a rider-day, the rule version, and the total the text requires; the test asserts equality and nothing else. That is an : intended behavior stated so a machine can check it and the rule's owner can review it.

The cost is permanent and larger than it looks. Ola joins the release path, so a fare change needs a sign-off measured in days — and Ola now maintains every example row forever, including edge cases engineering used to own. The rule set becomes a production dependency: when it cannot be loaded FareBridge declines to price, and a fare platform that cannot price is a product outage. categorySource is a path into FareBridge's own data model sitting in a document a non-engineer signs, so renaming that field needs a published vocabulary and a migration, not a search and replace. And because riders dispute fares months later, both rule versions stay computable for as long as re-pricing can reach back — a compatibility window somebody must eventually decommission.

One risk survives all of that: if the rule set is wrong, engine and tests agree on the same wrong answer. The rows only help while Ola authors them from the rule text rather than generating them from the engine, and if policy cannot sustain that load the design degrades into the self-confirming loop it was built to prevent. Correction A does not remove the need for someone outside engineering to be right. It makes visible who that is.

Correction B: make the recorded expectations explain themselves​

The second design kept the golden days and took away their authority.

First, each recorded expectation carried its reason:

before -> after, for RD-0431

- capCategory: "STANDARD", capBasisMinor: 1000, totalChargedMinor: 850
+ capCategory: "CONCESSION", capBasisMinor: 1150, totalChargedMinor: 350

That diff explains itself. Second, the team added assertions no acceptance keystroke can rewrite, stated over generated inputs rather than recorded ones. A property is a claim of the form for all inputs satisfying a precondition, this predicate holds:

fc.assert(fc.property(arbitraryRiderDay(), (day) => {
const priced = priceServiceDay(day, {ruleSetAsOf: '2026-07-01'});
expect(priced.capCategory).toBe(day.entitlement.categoryAtServiceDayEnd);
}));

Third, the gate stopped accepting undeclared change. A commit rewriting recorded expectations had to name the published rule change it implemented, and CAP-4 v2 predicted movement only for rider-days at the cap once surcharges counted — not a change of category. The four lapsed-concession expectations would have been refused.

Ask what authorises a release here, because it is not the gate. Under Correction A the sign-off is Ola's; under B it is the declaration, and the gate's only job is to refuse whatever the declaration did not predict. That is also B's weakness. The property at its heart states a policy fact — which record the category comes from — and it is maintained by the people whose code it checks, so it is rewritable for exactly the reason a recorded expectation is rewritable by -u. Under A the same fact lives in categorySource; run both and the two copies can disagree.

Nor is B free in the ordinary ways. Wider expectations churn on every internal rename, so sixty files move when nothing behavioral has. Generators need domain constraints, and maintaining them is real work. Attribution adds a step to every change. And the corpus still proves only that a change was declared, never that the declared rule is right.

One guard belongs to neither correction and is worth having under either: distinct types for the observed and the verified category, so passing one where the other belongs no longer compiles. That would have blocked this substitution. It would not have made a recorded total able to state a rule.

Concepts in this story4 concepts
Characterization testSoftware evolution · Test intent

A test whose expected value was recorded from the system’s own observed output so that later change becomes visible. It documents what the code currently does, which is not the same as what the code should do: a pass establishes agreement with an earlier observation, not conformance to a rule.

Expectation provenanceSoftware evolution · Test intent

Where a test’s expected value came from — a stated rule, policy, or invariant, or an observation of the system itself. Provenance decides what a passing test can prove, and an expectation recorded from output cannot separate a change made for the intended reason from one made for an unintended reason.

Executable specificationSoftware evolution · Specification

A statement of intended behavior written so that a machine can check it and the rule’s owner can review it — an invariant over all inputs, or example rows authored from the rule text rather than generated from the running system. Its value depends on being authored independently of the code it checks.

Effective-dated ruleSoftware evolution · Behavior versioning

A business rule published with a version and a validity window, so behavior can change on a date without erasing what was correct before. Effective dating makes a behavior change reviewable and reversible, and it obliges the system to keep older versions computable for as long as re-pricing, disputes, or audits can reach back.

What Kestrelway measured separately​

Hugo scored one build twice — the build that shipped on 29 June, once with the old checks and once with the new ones — over the sixty golden days, a clearly fictional 12,000 generated rider-days from a generator that respects Kestrelway's one-directional staleness rule, and the forty-eight example rows Ola authored. Every row below is something a check reported about the defective build, not behavior that survived a correction. The figures were constructed for this story.

What a check reported about the 29 June buildResult
Old check — recorded rider-day totals reproduced60 / 60
Old check — "never charged above the cap" violations0 / 12,000
New check — cap selected from a non-authoritative record214 / 12,000
New check — policy-authored example rows failing1 / 48
Charges submitted after a failed gate0

The two old checks were green during the incident, which is why they keep their own lines: a dashboard reporting only "regression suite passing" and "cap never exceeded" would have looked healthy for all eleven service days.

The two new checks are the corrections working. The 214 are generated rider-days where the category assertion refused a release, and the one failing example row is Ola's lapsed-concession row disagreeing with the engine — exactly what an independently authored row is for. Separately, of the 23 recorded expectations that changed on 29 June, the attribution gate would have refused 4 as movement no declared rule predicted.

Loading the corrected production topology…

The two designs disagree about where intent should be written — in an artifact the rule's owner reviews, or in assertions engineering maintains. They agree about the boundary: the expectation must come from somewhere the code cannot rewrite, and the release gate may only refuse. It may never authorize.

Transfer the provenance question​

A support-automation team runs a 900-question regression suite for its assistant. The expected answers were recorded from the previous model release and approved by the team. A customer then changes a refund window from 30 days to 14. Retrieval is updated, the new release answers "14 days", and forty-one questions fail.

Before touching the failures, ask three questions:

  1. What did the green suite establish, in the weeks before the change?
  2. Which artifact should own the expected answer, if not the previous release's output?
  3. Which assertion would fail if the new answer were right for the wrong reason?

"The suite was green" answers none of them, and neither does "we re-approved the answers". The same three questions work on a report whose expected figures were copied from last quarter's run, on a permissions matrix exported from production, and on a pricing table captured from the pricing service itself.

The rule Hugo kept was short:

A recorded expectation proves that behavior has not changed. Only a stated rule proves that behavior is the one you meant.

Evidence and fiction note​

The mechanics of recorded-output checks are documented by the tools themselves. Jest's snapshot testing guide covers first-run creation, the update flags, the continuous-integration position, the instruction that the artifact be "committed alongside code changes, and reviewed as part of your code review process", and the warning quoted earlier. Vitest's snapshot guide states the scope of the guarantee: snapshot tests exist to make sure "the output of your functions does not change unexpectedly". Neither says a passing snapshot establishes correctness, and neither is criticised here. ApprovalTests describes the same family, "also known as Golden Master Tests or Snapshot Testing".

Describing what code does rather than stating what it should do is Michael Feathers's characterization test, from Working Effectively with Legacy Code (2004), paraphrased here rather than quoted. Pact's guidance on contract tests versus functional tests supports the wider point: a green compatibility check can be right about the messages crossing a boundary while establishing nothing about the business logic behind them — a contract check over Kestrelway's settlement payload would have been green throughout. The property definition is fast-check's own. The cost of keeping two rule versions computable is the migrate step of parallel change: "the supplier has to support two different versions".

Everything else is fictional: Kestrelway Transit, FareBridge, Ola Fadeyi, Hugo Bertrand, Grace Mbeki, the golden days, CAP-4, both rule sets, and every fare, cap, card product, entitlement, identifier, date, figure and replay result above, including the 2,610 rider-days and the 831,400 minor units. Nothing here describes a real transit operator, payment network, or jurisdiction, and none of it is guidance about fares, payments, or regulation. It is an engineering story about where a test's expected value comes from.