The replay that reproduced every event and changed the bill
Larkspur Energy is a fictional retail electricity supplier. It sells Larkspur Flex, a plan that prices consumption by time-of-use band and adds a demand charge derived from the largest half-hour a site draws inside a peak band. Its customers are homes and small businesses. It does not own meters: a fictional metering agent, Gridpost Metering, reads them and publishes the data. It does not own the market either — it submits a daily settlement position to a regional market operator and lives with what the operator settles.
Sanne Vermeer (streaming platform engineer) owns the two stream jobs that turn meter reads into money, and the tooling that replays them. Dov Aharoni (settlement analyst) owns the daily market submission and authorises backfills when metering data goes missing. Marta Oyelaran (billing operations lead) owns customer notices, corrections, and the policy that decides which plan a site is eligible for. Marta owns the consequence of this incident and was not part of the repair that caused it.
How a half-hour of electricity becomes a bill
Larkspur Flex is defined by five rules. All five are Larkspur's own, or its market's; none describes a real market.
| Rule | Value |
|---|---|
| Rate classes | OFF_PEAK, SHOULDER, PEAK, CRITICAL_PEAK |
| Band adjustment applied to each rated interval when computing the demand metric | 1.01 / 1.02 / 1.06 / 1.09 respectively |
| Coincident peak | Per half-hour, sum the site's adjusted kWh across its meter points; take the maximum over half-hours classed PEAK or CRITICAL_PEAK; double it to express kW |
| Plan eligibility | A rolling twelve-month coincident peak above 25 kW moves the site off Flex onto a demand tariff, with 14 days notice |
| Market submission | Trading day T is submitted on T+1, and the revision window closes on the 15th of the following month |
A site can have several meter points — a main supply, a dedicated circuit, a charger — and the coincident peak is a property of the site, computed across all of them. That detail is the whole story, and it looks like nothing.
Concepts in this story5 concepts
Resolving a value that lives outside the event using a stated time rather than whatever the value happens to be now. The stated time can mean two different things — the time the event describes, or a declared cut in when the system asserted the value — and a design that supplies only the first still returns today’s assertion about a past day.
The property that re-running a computation over the same input produces the same output, where the input includes every external value the computation reads. Reproducing offsets, keys, order, event timestamps, watermark progression, and code version is necessary and not sufficient; anything resolved at processing time is different on every run by construction.
Valid time is the period during which a fact is true of the world. Assertion time is the period during which a system held that description of it. A source recording only the current description can be correct at two different moments and give two different answers, and a rebuild that does not say which assertion it wants silently takes the newest one.
A derived value or constraint whose scope spans more than one ordering key. Per-key ordering and per-key replay scope are both defined below its granularity, so neither preserves it: a rebuild covering only some of the invariant’s keys can produce a result that no consistent rebuild of any single version would produce.
The point after which a derived record stops being a computation and becomes a commitment. Before it, a change is a correction; after it, the same change is a restatement that needs a scope, an owner, and an authorisation. Finality is a domain rule rather than a system property, and different boundaries have different owners and different reopening costs — some belong to another party and cannot be reopened at all.
The part Sanne's team got right
Meter reads land on a log topic, meter.reads.v4, keyed by meter point. A stream job called rater turns each raw read into a rated interval: the same consumption, plus the contract that governed it and the rate class that applied to it.
Resolving the contract is where most teams go wrong, and Sanne's team did not. A customer's contract changes — they switch plans, they move, a discount ends — so reading "the customer's current contract" would price a February interval with a July plan. Instead, rater resolves the contract against a versioned table: a table with a primary key and a time attribute, which keeps the period during which each value was valid for that key. The job asks for the contract as of the interval's own start time.
That is : resolving a value that lives outside the event using a stated time, rather than using whatever the value happens to be now. Larkspur had built it, tested it, and relied on it for two years.
Eight lines
-- rater, revision r14 (fictional job revision; Apache Flink SQL)
INSERT INTO rated_intervals
SELECT r.meter_point_id, r.site_id, r.interval_start, r.kwh,
c.contract_id, c.plan_code, cal.rate_class
FROM meter_reads AS r
JOIN contract_versions FOR SYSTEM_TIME AS OF r.interval_start AS c
ON r.meter_point_id = c.meter_point_id
JOIN rate_calendar FOR SYSTEM_TIME AS OF r.proc_time AS cal
ON cal.calendar_date = CAST(r.interval_start AS DATE)
AND cal.half_hour_index = HALF_HOUR_OF_DAY(r.interval_start);
HALF_HOUR_OF_DAY is a fictional helper function, and the two joins are shown in one statement to put the difference side by side. Everything else is ordinary.
The second join resolves the rate calendar: which half-hours of which days are off-peak, shoulder, peak, or critical peak, including holiday and market-declared exceptions. Four thousand rows in an operational store, changing a few times a year. rater reads it live on every interval, with no cache in front of it — the freshest possible read.
The two joins are written with the same keywords. They differ in one identifier. r.interval_start is an event-time attribute: when the described thing happened, carried in the record, unchanged by anything the pipeline does later. r.proc_time is a processing-time attribute: when the stage ran. It is different on every run, by construction — that is what makes it processing time.
So the contract dimension reproduces — as long as that table still holds the June version, which is a retention promise with a cost, and one Larkspur happened to be paying. The calendar dimension reproduces the value that the calendar holds at the moment the join executes.
Nobody noticed, because nothing had gone wrong yet.
A concentrator, and a repair nobody would refuse
On 2026-06-03 a Gridpost concentrator at a substation lost its backhaul. A concentrator is the box that collects readings from many meters and forwards them, so nothing stopped happening at the meters: they went on recording half-hourly, locally, exactly as designed. What stopped was publication. Reads for 1,842 meter points, spread across 1,206 sites, stopped arriving.
The reads came back on 2026-06-29 — twenty-six days of them, the oldest twenty-six days late. rater allows six hours of lateness — a record whose event time falls behind the job's watermark is late, and a watermark is only an assertion about how far event time has progressed, an estimate of completeness rather than a guarantee of it. What happens to a record past the allowance is configuration, and Larkspur's configuration routed it to a side stream.
June invoices issued on 2026-07-04. Affected sites were billed on estimates and flagged for re-billing once the parked reads were processed. That part was legitimate and expected. Estimated bills exist so missing data does not stop a billing cycle; re-billing them on actuals is the promise, not the violation.
Reprocessing a side stream needed a manifest and Gridpost's data-quality sign-off on the recovered reads, which arrived on 2026-07-19. Nobody connected that date to the market's revision window: the two calendars are owned by different teams and live in different systems.
On 2026-07-20, Dov authorised the backfill. Sanne replayed meter.reads.v4 from the recorded offsets for exactly the 1,842 affected meter points, using the same job artifact and the same revision r14.
Afterwards, Sanne checked the replay. Here is what was checked, and what was not.
| Checked, and true | Not checked |
|---|---|
| Offsets matched the manifest exactly, record for record | Which reference values the joins resolved |
| Records in equalled records out | Whether any resolved value had changed since the original run |
| No record was dropped as late; the side-stream count returned to zero | Whether the rated output meant the same thing |
| Checkpoints succeeded; the job artifact and revision were identical | Whether the replay's scope covered everything the downstream aggregate needed |
| Per-key order, event timestamps, and watermark progression matched | — |
Every row on the left is a real property, and every one of them held. Together they are not , because determinism is a property of the whole computation, and the log was only one of its inputs.
The calendar was never wrong
On 2026-06-18, ten days after the fact, the market operator gazetted 2026-06-08 — published a binding notice designating it — as a grid flex day: on a flex day, 17:00 to 20:00 is re-designated CRITICAL_PEAK. Retroactive designation is normal in Larkspur's fictional market; the operator decides after the event whether the event happened.
On 2026-06-19, Larkspur's tariff reference team applied it. One statement, the obviously correct business action:
UPDATE rate_calendar
SET rate_class = 'CRITICAL_PEAK'
WHERE calendar_date = '2026-06-08'
AND half_hour_index BETWEEN 34 AND 39;
From that moment, no query could ask what the calendar said on 2026-06-15. The table carried an updated_at column, which records when a row last changed — not what it previously said.
This is the distinction the incident turns on. The rate calendar was correct on 2026-06-08, when it said the evening was shoulder. It was correct on 2026-07-20, when it said the evening was critical peak. Those are two different correct answers, because they answer two different questions. One is : the period a fact is true of the world, against the period the system held that description.
So when the replay ran on 2026-07-20 and resolved the calendar as of its own processing time, it got the right answer to the wrong question. The evening intervals of 2026-06-08 came out CRITICAL_PEAK, with a 1.09 multiplier instead of 1.02.
Two things follow that are easy to get backwards.
First, a cache would not have helped; a fresher one would have helped least of all. The read was already live — freshness was the failure.
Second — and this is the part that survives one round of fixing — against a calendar that keeps one timeline, as Larkspur's does, resolving as of interval_start returns the July answer too. Asking a single-timeline table about 2026-06-08 asks what it now says about that day; there is no assertion history to consult, so there is nothing for an event-time join to find. Making it answer the other question is not a change of identifier but a change of table.
The replay never asked the only question that mattered: reproduce the meaning this period already had, or deliberately restate it with what we know now? Nobody asked, because nobody had a way to express either.
Site 4471, and a number from nowhere
SITE-4471 is a fictional bakery with three meter points. MP-88120 (main supply) and MP-88121 (oven circuit) sat behind the failed concentrator. MP-88122 (a charger circuit) reported through a different concentrator and arrived on time, every half-hour of 2026-06-08, as usual.
Three half-hours of that day matter. Consumption in kWh:
| Half-hour | Class in June | Class in July | MP-88120 | MP-88121 | MP-88122 |
|---|---|---|---|---|---|
| 08:00 | PEAK | PEAK | 4.90 | 3.20 | 0.40 |
| 17:30 | SHOULDER | CRITICAL_PEAK | 5.60 | 4.10 | 3.40 |
| 19:00 | SHOULDER | CRITICAL_PEAK | 4.20 | 3.00 | 3.90 |
The second stream job, site-rollup, is keyed by site and local date. When new rated intervals arrive for a site-day, it recomputes that site-day from its retained state plus the new rows, and writes one row per site and local date into settlement.site_daily — the table the market submission, the bill run and the eligibility policy all read. Its snapshots are immutable, so a report can pin one and cite an exact number; none of the three did. site-rollup's own state holds what it was given, rated the way it was rated.
And it has a tie-break: when a half-hour's rated intervals carry different rate classes, it takes the highest-precedence class present. That rule was added years earlier so a half-hour whose meter points had only partly arrived would not be silently dropped from a demand calculation. It had never been asked to adjudicate between two vocabularies.
Because the multiplier is applied to each rated interval by rater, using that row's own class, the site's evening total now depends on which rows were replayed.
| World | What it means | Peak-band maximum | Eligibility |
|---|---|---|---|
| Faithful reproduction | All three meter points rated under the calendar as asserted in June | 18.02 kW | Rolling max stays at its February value of 19.6 kW — stays on Flex |
| Deliberate restatement | All three rated under the calendar as asserted in July | 28.56 kW | Crosses 25 kW — migrate, as a decision someone made |
| What happened | Two meter points rated in July, one still carrying its June rating | 28.08 kW | Crosses 25 kW — migrated, with nobody deciding |
The arithmetic is short enough to check. Faithful: only 08:00 is peak-eligible, so (4.90 + 3.20 + 0.40) × 1.06 × 2 = 18.02. Restated: 17:30 becomes eligible and is the largest peak-band half-hour, so (5.60 + 4.10 + 3.40) × 1.09 × 2 = 28.56. Mixed: the two replayed rows carry 1.09 and the third still carries 1.02, so (5.60 × 1.09 + 4.10 × 1.09 + 3.40 × 1.02) × 2 = 28.08.
Each row was individually correct. The aggregate was incoherent not because a row was wrong but because a was computed across rows resolved under two different assertions. The site's coincident peak spans three meter-point keys, which the log Larkspur uses may place on one partition or on three; either way it orders each key only against itself. Nothing about that was broken: the log guarantees that a given key always lands in the same partition and that one partition is read in write order, and both held perfectly. The guarantee simply says nothing about a number computed across keys — and the replay manifest, scoped to the keys that had failed, was defined at the same granularity as the guarantee rather than at the granularity of the invariant.
One more property explains why nothing caught this. The replayed rows always carry the larger multiplier, and reclassification only ever promotes a half-hour into the peak band. So for any mixed site:
faithful ≤ mixed ≤ restated
18.02 28.08 28.56
The mixture can never exceed a consistent restatement or fall below a faithful reproduction, so every mixed value was plausible. No range check could fire, and the reconciliation report that compares Larkspur's own projection against the market's settled position showed the 2026-06-08 variance, aggregated across every site behind one market node, as 0.4% — under its 1% threshold.
A tie-break that skipped disagreeing half-hours would have produced 18.02 kW and hidden a real crossing instead of fabricating a false one — a different third state, not the absence of one.
The job that was working perfectly
Larkspur's plan-eligibility job runs monthly, on the 21st. On 2026-07-21 it read the settlement projection, computed each site's rolling twelve-month maximum coincident peak, found 41 sites above 25 kW, migrated them off Flex onto the demand tariff, and sent each one a statutory 14-day notice.
The job was correct. It read the projection it was designed to read and applied the policy it was designed to apply. Nobody in the repair knew it ran on the 21st, and it had no way to know a repair had happened.
The 1,206 sites split. For 817 of them, every affected meter point sat behind the failed concentrator, so 2026-06-08 was re-rated consistently under the July assertion — on a settled period, with nobody deciding. The 389 mixed sites add the scope failure on top of that. The 41 migrations span both groups.
| Measure | Value |
|---|---|
| Sites whose 2026-06-08 site-day was rated under two assertions | 389 |
| Sites re-rated entirely under the July assertion | 817 |
| Sites migrated off Flex on 2026-07-21 | 41 |
| Migrated only because a retroactive change reached a settled period | 33 |
| Above 25 kW regardless of which calendar applied | 8 |
| Above 25 kW under a consistent restatement, left below it by the mixture | 9 |
On 2026-08-04, the day the migration took effect, the bakery's re-billed June statement arrived carrying a demand charge computed from 28.08 kW. They called Marta's team. That was fifteen days after the repair, in a system nobody in the repair had touched, reported by a customer.
Two boundaries, one of them closed
The reason none of this could simply be recomputed away is : the point after which a derived record stops being a computation and becomes a commitment. Finality is a rule the business owns, not a property the pipeline has, and Larkspur had crossed two boundaries with two different owners.
The internal boundary was the eligibility decision. Larkspur could retract 41 notices — at the cost of a written retraction to each customer and a reportable notice error. Expensive, embarrassing, possible.
The external boundary was the market position for 2026-06-08. It was submitted on 2026-06-09, before the flex day was even gazetted. The revision window closed on 2026-07-15, five days before the backfill ran. After 2026-07-20, Larkspur's internal projection and the market's settled position for that day disagree permanently. Under this market's rules the settled position stands as the financial record, so Larkspur absorbs the difference and carries an unreconciled line into its own audit — and no amount of pipeline correctness reopens a window that belongs to somebody else.
Before 2026-07-04 and 2026-07-15 these changes would have been corrections. The backfill ran after both, so the same changes were restatements — and nothing in the pipeline knew the difference, because nothing in it knew a boundary existed.
Six shortcuts, and what each becomes:
- Resolve reference data as of processing time → give it an assertion history and resolve it as of a declared knowledge horizon.
- Update reference rows in place → append an assertion interval and keep what the row said before.
- Scope a replay by the keys that broke → close the scope over every key the invariant spans.
- Let a replay mean whatever it produces → declare, before it runs, whether it reproduces or restates.
- Treat a recomputation as free → treat a change past a finality boundary as a restatement with an owner.
- Verify a replay by counting records and offsets → verify it by comparing meaning against the horizon it declared.
Two corrections that pay differently
Sanne, Dov and Marta compared two designs. Both keep event-time windowing, per-key ordering, the tariff rules and the tie-break; design B adds one refusal in front of them.
Design A — versioned reference authority with a declared horizon
Reference data gets a second clock. The calendar becomes a versioned table keyed by the day it describes and the interval during which Larkspur asserted it, so an as-of read can find a past assertion.
Every replay then declares two things it never used to declare:
manifest {
mode: REPRODUCE | RESTATE
knowledge_horizon: <assertion-time cut>
scope: closed over every meter point of every affected site-day
authorisation: required for RESTATE, naming each finality boundary crossed
}
REPRODUCE reproduces June's classes exactly, so a repair cannot change meaning. RESTATE produces a new, labelled version of the period and needs someone to sign for it. Scope closure re-rates every meter point of every affected site-day under one horizon, so no aggregate mixes vocabularies.
It costs four things and a fifth that is not engineering at all:
- Reference-data producers lose in-place updates. An
UPDATEbecomes an append with an assertion interval, and every consumer grows a horizon argument. - As-of reads are accurate only within retained assertion history, so the longest permitted replay window becomes a retention requirement. A four-thousand-row calendar is trivial to keep for years. Contract, network-tariff and loss-factor history is not.
- Genuine reference corrections get slower. A wrongly gazetted holiday used to be one statement; now it needs a scope and an authorisation.
- Scope closure inflates the repair. Fixing 1,842 meter points means replaying every meter point that shares an affected site — more compute, a longer backfill, more contention with the live job.
- When the reference data belongs to another company, the cost stops being engineering and becomes an inter-company agreement to publish versions instead of overwriting them.
Design B — materialized decision records
Instead of making the reference data answerable, make the derived record self-describing. rater emits an immutable decision:
rated-interval.v3 {
meter_point_id, site_id, interval_start, kwh,
contract_version_id, calendar_version_id, rate_class, band_multiplier,
decision_id, decided_at, rater_revision, supersedes
}
Downstream stages consume decisions, never raw reads, so their replay is reproducible for the reason the log alone was not: every value the stage needs is inside the record, and nothing is left to resolve from outside it. Only rater mints decisions, and only for intervals that have none. A wrong decision is never overwritten — it is superseded by a new decision carrying a supersedes link — and every aggregate declares which closed decision set it consumed.
Under this design the incident becomes visible rather than impossible: the site-day would hold rows citing two different calendar_version_id values, and the declared-set rule refuses to aggregate across them. A refused site-day yields no new projection row, so the eligibility job reads the stale one until somebody clears the refusal inside the billing cycle — where design A's manifest simply fails to validate before the replay runs.
It costs five things, and the first is larger than the other four:
- It cannot reconstruct an assertion nobody kept.
MP-88120andMP-88121had no decision for 2026-06-08, soratermust resolve their class from the calendar as it stands — and design B never asked producers to keep assertion history. A period with missing keys can only be restated, visibly, never reproduced. Design A can runREPRODUCE; design B cannot. That is a difference in kind. - Records grow several times over, forty-eight per meter point per day. At Larkspur's fictional volume the topic grows from 118 GB to 412 GB a month.
- A second authoritative store appears, with its own retention obligation and its own capacity to drift from the first.
- Every consumer and every report must handle supersession and name a decision set, which makes reports harder to explain to people who do not work on the pipeline.
- Fixing a genuinely wrong reference value becomes harder than under design A. Design A declares one restatement and replays. Design B mints superseding decisions for every affected key, then re-aggregates a declared set.
The last cost is the one worth arguing about — the first is simply larger — because it inverts the intuition that the more explicit design is strictly safer. Design B is safer against silent change and worse at deliberate change.
The rule Sanne's team wrote down was narrow on purpose:
- Prefer A when the reference dimensions are few, small, and owned by teams whose write path you can change.
- Prefer B when they are numerous, owned elsewhere, or when an audit needs each derived record to explain itself without a time-travel query — but only if faithfully reproducing a closed period is not a requirement. When it is, B needs retained reference versions too, and if the owner will not publish them you must record each version as it arrives and pay for both.
Neither design reopens the market's revision window. Neither un-sends 41 notices. Both convert an invisible, unowned change in meaning into a visible, owned decision. Only A can also reproduce the meaning a closed period already had.
Transfer: the payslip that was consistent and still wrong
A clock-terminal outage means a week of shifts arrives eleven days late. In that interval, a backdated agreement raises the overtime multiplier with an effective date inside the gap. The pipeline replays the shifts in perfect order, same code, same keys. A weekly overtime threshold spans a worker's shifts across several cost centres — several keys — and only some of those shifts were in the replay scope. Payslips for that week were issued and paid.
Diagnose the variant. The team closes the replay scope over the whole worker-week, but does not declare a knowledge horizon. Which of the two failures does that fix, which does it leave, and what does the resulting payslip mean?
Then choose. The agreement is negotiated by a third party that publishes a single current rate sheet and will not publish versions. Which design applies, and what does it cost?
The first answer is the more useful one. Scope closure removes the incoherence — every shift in the week is now rated under one assertion — but with no declared horizon that assertion is "now," so the rebuild quietly restates a period that was already paid. The payslip becomes internally consistent and still is not the payslip that was owed. Consistency is not faithfulness.
An immutable ordered log makes the input reproducible. It says nothing about the values your computation reads from outside it, the moment it read them, or the scope it read them over. Decide which of those a rebuild is supposed to reproduce — and decide it before the rebuild runs, because afterwards the only question left is who signs for the restatement.
Concepts in this story5 concepts
Resolving a value that lives outside the event using a stated time rather than whatever the value happens to be now. The stated time can mean two different things — the time the event describes, or a declared cut in when the system asserted the value — and a design that supplies only the first still returns today’s assertion about a past day.
The property that re-running a computation over the same input produces the same output, where the input includes every external value the computation reads. Reproducing offsets, keys, order, event timestamps, watermark progression, and code version is necessary and not sufficient; anything resolved at processing time is different on every run by construction.
Valid time is the period during which a fact is true of the world. Assertion time is the period during which a system held that description of it. A source recording only the current description can be correct at two different moments and give two different answers, and a rebuild that does not say which assertion it wants silently takes the newest one.
A derived value or constraint whose scope spans more than one ordering key. Per-key ordering and per-key replay scope are both defined below its granularity, so neither preserves it: a rebuild covering only some of the invariant’s keys can produce a result that no consistent rebuild of any single version would produce.
The point after which a derived record stops being a computation and becomes a commitment. Before it, a change is a correction; after it, the same change is a restatement that needs a scope, an owner, and an authorisation. Finality is a domain rule rather than a system property, and different boundaries have different owners and different reopening costs — some belong to another party and cannot be reopened at all.
Evidence and fiction note. Larkspur Energy, Larkspur Flex, Gridpost Metering, the market operator, Sanne Vermeer, Dov Aharoni, Marta Oyelaran, every identifier, all tariff and market rules, the band adjustment factors, the 25 kW threshold, the dates, volumes, metrics, and every replay result above are fictional. Larkspur's six-hour lateness allowance and its side stream are that fictional pipeline's own configuration, not documented behaviour of any engine.
Technical grounding is bounded to mechanism claims. Apache Flink 2.3.0 SQL joins documents that an event-time temporal join can retrieve a versioned table's value "as it was at some point in the past," that a processing-time temporal join "will always return the most up-to-date value for a given key," and that "The result is not deterministic for processing-time." Flink's own determinism guidance lists a lookup join on an evolving source among the causes of non-deterministic updates — the hazard is documented, not folklore, and neither source describes a defect. Flink's versioned-tables concept page, quoted here from the 2.2.1 documentation because the page moved between releases, states that such a table needs a PRIMARY KEY and a time attribute and "maintain the period for which each value was valid for that key." Apache Kafka 4.3 documents that events with the same key go to the same partition and that a consumer reads a topic-partition in write order, and that events "can be read as often as needed" within configured retention — it makes no claim about ordering across keys. KIP-889 records that Kafka Streams reaches as-of lookups through a different mechanism, versioned state stores, "accurate results for calls to get(key, asOfTimestamp) where the provided timestamp bound is within history retention" — which is where the retention cost in design A comes from. The Apache Iceberg specification defines a snapshot as "the state of a table at some time" and states that data and metadata files "are immutable until they are deleted."
These sources are engine-specific in three places that matter: which temporal semantics a join keyword selects, how far back as-of history is retained, and what happens to a late record. None of them establishes Larkspur's market, tariff, thresholds, workload, metrics, or incident, and none should be read as describing a real energy market or metering agent. The event-time, processing-time, watermark and lateness vocabulary used here descends from Akidau et al., The Dataflow Model, PVLDB volume 8, 2015; no sentence of that paper is quoted.