Skip to main content

The interruption that stopped the sentence and not the payment

· 40 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about barge-in, separate cancellation domains, and why a spoken assistant must never describe an effect it does not own.

Ferrowick is a fictional construction-operations platform used by mid-sized building contractors. Customers run whole projects on it: purchase orders, delivery receipting, retention tracking, subcontractor valuations, and the weekly supplier payment run. Ferrowick holds no customer money; payments are initiated through Norwell Pay, a fictional licensed payment provider that submits into a bank scheme.

Kelter is Ferrowick's realtime site assistant, and it exists because of one physical fact. The person who knows what actually arrived is standing in a delivery bay in gloves and hearing protection, beside a reversing lorry, with a clipboard in one hand. She is not going to open a laptop. Kelter runs on a site handset or a paired headset, connects over WebRTC to a regional media edge, listens, and speaks.

A payment passes three decisions, made by different people at different times. On Monday the project manager approves the week's run in the web application, deciding what may be paid at all. During the week the site operations lead receipts deliveries, deciding what arrived. When retention conditions are met she releases the matching payment by voice, deciding what is sent now and in what order.

That third decision matters, because a project's available balance is finite and sequenced. Releasing early does not move money to the wrong week; it consumes the headroom a later payment depends on. A supplier who is not paid does not deliver, and a crew that arrives to an empty bay stands idle at full cost.

Priya Raghavan is Site Operations Lead on the fictional Halden Row project; she receipts deliveries and releases approved payments. Tomas Idrissi is a Staff Engineer who owns Ferrowick's realtime interaction plane — media session, connection epoch, turn state, playout, interruption handling. Dela Okonkwo is a Principal Engineer who owns the payment-effect plane — the gateway, the operation ledger, the Norwell Pay integration. Ruth Vance is Head of Payment Operations, and owns what a customer is told when the state of a payment is uncertain.

Tomas and Dela both built correct systems. The sentence that caused this incident lived in Tomas's plane and was about Dela's, and nobody had ever decided which of them owned it.

How a supplier payment normally leaves Ferrowick​

Loading the Ferrowick payment journey…
Concepts in this story5 concepts
Cancellation domainRealtime multimodal · Interruption and cancellation

One region of state that a single cancellation actually reaches, together with the system that owns it. A live interaction has several — audio playout, model generation, turn and command intent, the tool request, and the external effect — and each acknowledges separately, at a different speed, with a different meaning. How fast an acknowledgement arrives tells you nothing about what it proves, and the domain that owns the effect may never acknowledge at all.

Committed turnRealtime multimodal · Turn-taking

The point at which a system decides a person has finished speaking and treats their words as an input it will act on. Silence thresholds and semantic endpointers make that decision probabilistic, so a committed turn is a guess about a human thought, not a fact about it — and an irreversible effect must not depend on one.

Unheard-output cursorRealtime multimodal · Session continuity

The boundary between output a system generated or sent and output a person actually heard. Realtime models emit audio faster than realtime, so an interruption leaves buffered speech that was delivered and never played. Conversation history must be reconciled to the played position, or a resumed session will confidently refer to a confirmation nobody received.

Irreversible commit pointDistributed systems · Effect safety

The moment in an effect’s lifecycle after which cancellation no longer exists, and the only remaining option is a separate compensating request that another party may refuse. Before it, cancellation can be exact because no external system has been told anything; after it, a cancellation is a negotiation. Holding an effect moves this point later; it never removes it.

Stale-result suppressionRealtime multimodal · Conversational tools

Dropping a late asynchronous result from the current conversation because its turn, topic, or command has moved on. It is the right conversational policy and the wrong records policy: a receipt must always enter the authoritative effect log, while whether it is spoken aloud is a separate decision with a separate owner.

Five groups own different things along that path. The handset captures audio and plays replies through a local playout queue. A regional media edge terminates the WebRTC connection, and a stateful session owner holds the connection epoch, the turn state, and a conversation checkpoint. A realtime speech model does three jobs: understand supplier names and amounts under site noise, decide when Priya has finished a thought, and turn "release the Marlstone invoice" into a typed command resolved against the approved run. Ferrowick's payment-effect gateway records a ReleaseCommand in an operation ledger and asks Norwell Pay to submit it into the bank scheme.

The model has real work and no ownership: not the approved run, the balance, the command state machine, the decision to submit, or the record of what happened. A better model would misunderstand fewer supplier names and misjudge fewer pauses. It would not change who owns anything.

The design the team trusted​

Four decisions, each individually defensible, all in one code path.

The payment tool was non-blocking. A blocking submission produced four to nine seconds of dead air while the gateway waited on Norwell Pay; operators talked into the silence and turn state degraded. So Kelter dispatches the release and keeps talking. Platforms support this directly: Google's Live API documents a behavior field with a "NON_BLOCKING" value and a scheduling parameter with three modes — "INTERRUPT", "WHEN_IDLE", "SILENT" — so a slow tool cannot freeze a conversation.

Barge-in was handled carefully. Barge-in — a person speaking over the assistant while it is still talking — is only detectable because Kelter listens while it speaks, which is what full duplex means. Tomas built that path and measured it. On speech start, Kelter cancels model generation, clears the client's playout queue, marks the turn abandoned, and sends a cancellation to any tool call that turn started. It was fast, consistent, and demonstrated beautifully.

Late results were suppressed. A background result must not barge into a newer topic, so a result whose command was cancelled, or whose turn is gone, is dropped. Anyone who has watched a stale tool result interrupt a live conversation has written this rule.

Release carried no confirmation step. This is the decision nobody wrote down as a decision, and it needs saying, because moving money on one unconfirmed utterance sounds careless. Ferrowick's reasoning was that authorisation had already happened upstream — run approved on Monday, delivery receipted — so what remained was a sequencing choice, and operators already found Kelter slow. A team that has authorised the payee and amount upstream, under pressure to reduce friction downstream, plausibly treats release as low-ceremony work. That classification is defensible. It is also the one the incident falsifies.

The mental model survived design review:

Interruption is handled end to end. We cancel generation, we cancel audio, we abandon the turn, and we cancel the tool. Nothing in flight survives an interruption.

No missing timeout, no unhandled error, no model holding a credential it should not have.

One sentence, two turns​

A turn does not end when a person stops talking. It ends when the system decides they have finished — and every mechanism for deciding is a rule applied to silence or to words, never a fact about a thought.

Both platforms expose the dial and describe the trade-off honestly. OpenAI's Realtime API offers server voice-activity detection — silence-based chunking with threshold, prefix_padding_ms, silence_duration_ms — or semantic detection, which uses "a semantic classifier to detect when the user has finished speaking, based on the words they have uttered", with an eagerness property of low, medium, high, or auto. Low eagerness lets the user take their time; high eagerness feels quick. Google's Live API exposes equivalent knobs under automaticActivityDetection, warning that a silenceDurationMs under roughly 200 ms risks splitting an utterance into fragments, with 500–800 ms a reasonable balance.

Here is what Priya said:

"Release the Marlstone invoice — actually no, hold on, the retention isn't released yet."

She stopped after "invoice" to check a delivery note. Ferrowick's endpointer was configured to commit after roughly 340 ms of silence, and it committed — but so would any documented setting, including the 500–800 ms band Google calls balanced, because she was silent for a second and a half. Her correction arrived when it did not because she was cut off mid-breath, but because she heard Kelter begin confirming something she had not finished asking for.

That is the uncomfortable version, and the one the numbers support: no endpointing configuration would have prevented this. A more patient dial buys a longer window, and every window closes while somebody is still thinking. Half an utterance had become a command, her correction had become a turn that did not exist yet, and what separated them was an ordinary second and a half of human silence.

Wednesday at 09:11:58​

All offsets are fictional, measured from a single reference of 09:11:58.000 and chosen to be consistent with documented endpointing ranges and ordinary network latencies. They are not measurements and not a provider's performance claim.

The turn committed at +1.480. At +2.100 the orchestrator dispatched the release as a non-blocking tool call and Kelter began speaking a confirmation; at +2.400 the gateway posted it to Norwell Pay. Priya's correction landed at +2.720. Generation stopped, the playout queue cleared at +2.760 after 620 ms of audio had played, the turn was abandoned, and a cancellation went to the tool call — arriving at +2.900, after Norwell Pay had already forwarded the payment. The scheme accepted at +4.100. The receipt returned to the conversation at +4.380 and was suppressed. At +5.000 Kelter said: "Held — nothing was sent."

Then a routine reconnection at +223 s, a refused duplicate submission at +364 s, and at +2,469 s a treasury balance alert that blocked the concrete supplier's payment. Forty-one fictional minutes after the money left, the conversation, the turn state, and Kelter all still agreed it had not. The Friday pour slipped, and on Monday a groundworks crew stood in an empty bay at full cost.

Loading the incident sequence…

What each acknowledgement actually proves​

The code said cancel(). What happened was five independent operations against five owners.

DomainOwnerMechanismLatency (fictional)Its acknowledgement provesIt does not prove
Playoutthe handsetclear the local audio queue~40 msOutput stopped reaching her earAnything about the model, the tool, or the money
Generationrealtime providerprovider cancels the response~20 msOutput stopped being producedThat output already sent is gone from history
Turn and command intentFerrowick orchestratormark the turn abandoned~60 msThe application will not act on that turn againThat anything already dispatched stopped
Tool requestthe tool transporta cancellation notification, or a closed stream~120 ms to sendOnly that a request to stop was sentReceipt, honour, or timeliness
External effectNorwell Pay, then the schemenone once forwarded—Whatever the ledger's receipts sayThat anything stopped, unless a receipt says so

Four of those five acknowledged. One never did, and it was the one holding the money.

Now read the last column and notice it has nothing to do with the column before it. Speed tells you nothing about proof: the two fastest acknowledgements are genuinely informative, and the slowest of the four proves least, because a sent cancellation is only a sent cancellation. Any code path treating any of them as an answer about the effect is confidently wrong at a low, non-zero rate.

The specifications are direct about the fourth row. The Model Context Protocol revision dated 2026-07-28 defines cancellation as a request — a notifications/cancelled notification over stdio, or a closed response stream over Streamable HTTP — then states that "due to network latency, cancellation notifications may arrive after request processing has completed, and potentially after a response has already been sent", and that servers may ignore one if the request is unknown, if processing has completed, or if it cannot be cancelled. A2A 1.0.0 is equally blunt about remote tasks: "the server will attempt to cancel the task, but success is not guaranteed", and a task already terminal must fail with TaskNotCancelableError. The Streamable HTTP rule has a sharper edge: closing the stream is the signal, and a server must treat a client disconnect as cancellation — so on a site with an intermittent uplink, an intentional cancellation and a lost signal are indistinguishable at the transport.

Telephony settled this decades ago. In SIP, CANCEL asks a server to stop processing a pending request; once a final 2xx has been sent the session exists and you tear it down with BYE. Voice interfaces only shortened the pending window.

"Held — nothing was sent."​

At +5.000 Kelter said something false in a calm, confident voice. Not lying, not hallucinating — reading the plane it had. Turn state said the command was cancelled, because the orchestrator marked it so when it sent the cancellation. There was nowhere to record the only true statement available: we asked, and we do not yet know. Priya heard "held", believed it, and moved to the next delivery.

At +4.100 the scheme accepted. In Ferrowick's fictional contract that transition is the irreversible commit point: after it there is no cancellation of any kind, only a recall request, which is a request the receiving institution may refuse.

The lifecycle Ferrowick shipped afterwards says this structurally rather than in a comment:

PROPOSED ──confirm──▶ CONFIRMED ──hold──▶ HELD ──release──▶ SUBMITTING
└──────────submit──────────────────▶ SUBMITTING

PROPOSED | CONFIRMED | HELD ──cancel──▶ CANCELLED_PRE_SUBMIT (exact)

SUBMITTING ──cancel──▶ CANCEL_RACE_PENDING
CANCEL_RACE_PENDING ──▶ CANCELLED_PRE_SUBMIT | SUBMITTED_UNCERTAIN
| ACCEPTED_BY_SCHEME | REJECTED

SUBMITTED_UNCERTAIN ──reconcile──▶ ACCEPTED_BY_SCHEME | REJECTED
ACCEPTED_BY_SCHEME ──recall──▶ RECALL_REQUESTED ──▶ RECALL_HONOURED | RECALL_REFUSED

The same contract in words, because a picture built from box-drawing characters is not an equivalent for every reader. A release is cancellable exactly and only while PROPOSED, CONFIRMED, or HELD. Once SUBMITTING, a cancellation moves it to CANCEL_RACE_PENDING, which resolves four ways — cancelled before submission, uncertain, accepted, or rejected — and Ferrowick cannot tell which until a receipt says so. SUBMITTED_UNCERTAIN reconciles later to accepted or rejected. ACCEPTED_BY_SCHEME leads onward only to settlement or to a refusable recall request. The diagram omits terminal accounting states.

CANCEL_RACE_PENDING is the state the original design lacked, and its absence is the defect in one line: without it, the only way to record "we asked it to stop" was to record "it stopped".

Dela wrote the boundary of cancellable at the top of the incident document:

A release is cancellable exactly while no external party has been told about it. Everything after that is a race, an uncertainty, or a negotiation.

The receipt that was thrown away​

At +4.380 the accepted receipt came back as a late result, and suppression dropped it, because its command was marked cancelled.

This is the part that stings, because the rule is right. Every documented scheduling mode exists to control exactly this: INTERRUPT when the result matters more than the current utterance, WHEN_IDLE when it can wait, SILENT when the model should simply know it. MCP goes further and tells clients to "ignore any response to the cancelled request that arrives afterward" — sound advice about a request/response pair, catastrophic advice about a receipt.

Two decisions had been collapsed into one:

DecisionCorrect ownerCorrect default
Does this late result enter the authoritative record?the payment planeAlways. A receipt is evidence, and evidence is not conversational.
Does this late result enter the current conversation?the interaction planeIt depends: interrupt, wait for idle, stay silent, or drop.

One rule, owned by the interaction team, was doing both jobs. The ledger did receive the receipt — Dela's plane was correct throughout — and a payment-status view in the web application would have shown it to anyone who looked. Nobody looked, because every surface that reached Priya where she was standing derived from the conversation, and nothing paged anyone. The evidence that would have corrected her existed, in the right place, and never travelled to the delivery bay.

The sentence nobody heard​

Three lifetimes are in play, and separating them matters before a handover breaks one. A turn lives inside a connection; a connection is capped and replaced by a new epoch when re-established; a session can outlive both, restored through a resumption handle. Three lifetimes, three owners, and none owns the effect.

At +223 s an LTE handover dropped the connection. This is not exotic: Google's Live API documents a connection limit of "around 10 minutes", a GoAway message with a timeLeft field, and resumption handles valid for two hours — so on that platform a one-hour conversation reconnects several times by design.

Kelter resumed, and minutes later said something worse than "held": it referred to "the confirmation I gave you at 09:12". It had never given that confirmation. It had generated it.

Three cursors existed at +2.760, and no two agreed:

CursorHeld byValue
Generatedthe providerthe complete confirmation
Sent to the clientthe providerseveral seconds of audio and its transcript
Played to the personthe handset620 ms

The gap between the last two is not a bug. Realtime models emit audio faster than realtime, so the client holds speech that has not been played — which is exactly why truncation exists. OpenAI's conversation guide tells the client to "note how much of the last audio response was played before the interruption", then send a conversation.item.truncate carrying that position "to remove the unplayed portion of the model's last response from the conversation".

Here the platforms diverge in a way no abstraction hides:

  • Google reconciles server-side to what it sent. On interruption "the ongoing generation is canceled and discarded. Only the information already sent to the client is retained in the session history." The client is separately told it "should stop playing audio and clear queued playback" — a media action with no history consequence.
  • OpenAI reconciles to what was played, but only if the client does the work. The server cancels the response and emits response.cancelled; the client must send the truncation event. The documentation warns that "the realtime model doesn't have enough information to precisely align transcript and audio", so the alignment is approximate even when you do it.

Ferrowick cleared the playout queue and never reported the position upward. No truncation event was sent, so its checkpoint stored the assistant's turn as sent, not as heard, and the resumed history held two contradictory assistant turns: a full confirmation of a submission, and a later statement that nothing was sent.

Asking a model to resolve a contradiction inside its own history is not a control. The control is to stop writing the contradiction: report the playout position and truncate history to it, accepting a bounded claim in exchange. History then agrees with what was played, at the resolution the client reports — which does not establish that anyone perceived it. Audio played into hearing protection beside a running compressor is played, not necessarily heard.

What the ledger could say, and what it could not​

Ferrowick already had most of the machinery, which is why the postmortem was uncomfortable. The operation ledger was correct — every transition appended with a receipt and a clock reading. Nobody read it before speaking.

Idempotent submission was correct and irrelevant to the divergence. At +364 s, still believing nothing had been sent, Priya asked Kelter to send the Marlstone payment "properly now". The duplicate window refused the second submission on the same commandId — exactly as designed, preventing a real duplicate — and did nothing about the belief that the first had not happened. The claim needs its boundary attached: at most one accepted scheme submission per commandId, at the gateway boundary, within the duplicate window, with SUBMITTED_UNCERTAIN as a real state reconciliation must resolve. That is not a guarantee across Norwell Pay or the scheme.

Fencing would not have helped. A fencing token prevents a new stale write at the boundary that enforces the token. Norwell Pay does not enforce Ferrowick's token, and by +2.900 the submission had been forwarded. Fencing never retracts a call already sent.

The ledger's own authority has a boundary. It is authoritative for what Ferrowick has observed. Norwell Pay's status can lag, so SUBMITTED_UNCERTAIN is not an edge state to engineer away; it is a normal answer Kelter must be able to say out loud.

And there is a version of Wednesday morning in which none of this happened. The gateway posted at +2.400 and the cancellation arrived at +2.900; 400 ms earlier it would have landed at +2.500, before Norwell Pay forwarded, and the command would have moved to CANCELLED_PRE_SUBMIT. "Held — nothing was sent" would have been true. That is why the design shipped: it was right most of the time. But even then the old code printed that sentence from turn state before it knew the outcome — correct by luck at a rate indistinguishable from correctness. A control whose correctness depends on winning a race is not a control.

How often does the race go the other way? Ferrowick could not say: a few hundred fictional releases a week, and four interruptions inside a submission window last quarter — enough to know the shape recurs, far too few to estimate how often it resolves badly. A rate you cannot estimate argues for the design that fails legibly over the one that fails rarely.

The release boundary in code​

Before, the handler cancelled everything it could reach and then described the world:

// Tempting: one cancel() over four owners, then narrate from turn state.
async function onSpeechStarted(session: Session) {
await provider.cancelResponse(session.activeResponseId); // generation
client.clearPlayoutQueue(); // media
const turn = turns.abandon(session.activeTurnId); // turn state
for (const call of turn.toolCalls) {
await tools.cancel(call.id); // a REQUEST, not a result
commands.markCancelled(call.commandId); // <-- the false statement
}
return speak('Held — nothing was sent.'); // the wrong plane speaking
}

function onToolResult(result: ToolResult) {
if (commands.isCancelled(result.commandId)) return; // receipt discarded
return conversation.deliver(result);
}

After, cancellation still cancels everything it can reach — and narration moves to the plane that owns the effect:

async function onSpeechStarted(session: Session) {
// Cancel only what this plane owns, and record each domain separately.
const stopped = {
generation: await provider.cancelResponse(session.activeResponseId),
playoutCursorMs: client.clearPlayoutQueueAndReportCursor(),
};

// Reconcile history to what was PLAYED, not to what was sent.
await session.truncateAssistantTurn(session.activeItemId, stopped.playoutCursorMs);
const turn = turns.abandon(session.activeTurnId);

// A cancellation is a request. Record the request, never the outcome.
for (const call of turn.toolCalls) {
await tools.requestCancel(call.id);
await ledger.recordCancelRequested(call.commandId, {
requestedAt: clock.now(), turnId: turn.id, epoch: session.epoch,
}); // -> CANCELLED_PRE_SUBMIT, or SUBMITTING -> CANCEL_RACE_PENDING
}

// Speak only what the plane that owns the effect currently reports.
const states = await ledger.currentStates(turn.toolCalls.map((c) => c.commandId));
return speak(narrateByDomain(stopped, states));
}

async function onToolResult(result: ToolResult) {
await ledger.append(result); // ALWAYS. A receipt is evidence.
divergence.check(result.commandId); // see Design B for the trigger

const schedule = conversationPolicy.scheduleFor(result); // interrupt|idle|silent|drop
if (schedule !== 'drop') conversation.deliver(result, schedule);
}

narrateByDomain composes one clause per domain, each from a state its speaker may report:

"I stopped speaking, and the provider acknowledged that it stopped generating. I asked our payment gateway to stop the release; the gateway's record says it had already been forwarded, and the scheme has accepted it. I have flagged this to payment operations, who decide whether to request a recall — and a recall is a request the receiving bank can refuse."

Notice what it does not say. It does not say Kelter asked the bank, because nobody did: the orchestrator asked a tool transport, and the effect domain has no cancellation mechanism once a payment is forwarded. It does not say a recall has been opened, because opening one is staffed human work in Ruth Vance's team. Nobody enjoys this sentence. It is the only true one available, and every clause names a plane its speaker owns or a record it has read.

The same change as an ordered contrast:

#Old shortcutCorrected path
1Clear playout, tell nobody the positionClear playout and report the playout cursor
2—Truncate the assistant's turn in history to that cursor
3Send a tool cancellation and record the command as cancelledSend a tool cancellation and record that a cancellation was requested
4Suppress any late result for a cancelled command, everywhereAppend every late result to the ledger; schedule conversation delivery separately
5Say "nothing was sent"Read the ledger and say what it says, including "I do not know yet"

Where the providers disagree​

An abstraction over "realtime voice" will hide these differences, and hiding them is how a team ends up with two truths about one conversation. Checked 2026-08-10.

QuestionOpenAI RealtimeGoogle Gemini Live APIMCP 2026-07-28A2A 1.0.0Bank scheme (fictional)
Who reconciles history to what was heard?The client, via conversation.item.truncate; alignment documented as impreciseThe server, to what was "already sent to the client"; the client clears playback———
Can an in-flight tool call be cancelled?Not documented on the realtime guidesNot documented; tool responses handled manually by the clientA request; servers may ignore it"Success is not guaranteed"; TaskNotCancelableErrorNo — only a refusable recall request
What happens to a late result?Not documented on the realtime guidesscheduling: INTERRUPT, WHEN_IDLE, SILENTClient should ignore a late responsePush delivery; unreachable clients undocumentedThe receipt is written to the ledger regardless
How long does the conversation last?A 60-minute maximum session duration; no separate connection cap stated on the guides reviewed"Without compression", a 15-minute audio session (2 minutes with video) — context-window compression removes that ceiling — inside a connection of "around 10 minutes", which it does not, plus a two-hour resumption handle———

The first row is a genuine architectural divergence: on one platform the client owns the truth of the transcript, on the other the server reconciles to a cursor the client does not control. The second is an absence rather than a difference — nothing reviewed here defines cancellation of a business effect a realtime tool call has already started, because no protocol can. That is not a documentation gap awaiting a release; it is a property of the world.

The fourth row is the one most likely to be flattened by an abstraction layer, and flattening it changes the architecture. Both platforms cap a conversation, and the caps are not the same shape — nor are they the same kind. One long session has a single ceiling, and reconnection is an exceptional event you handle when the network fails. On the other, two ceilings stack and only one of them is removable: context-window compression lifts the session limit, while the connection limit of around ten minutes stays. So the durable constraint is the one you cannot configure away, reconnection is a scheduled event several times an hour whether anything fails or not, and the resumed history in this story is a routine code path rather than a contrivance. "The session ended" is not one fact but two, and only one of them is yours to change. Google's tool documentation also states that asynchronous function calling "is not yet supported in Gemini 3.1 Flash Live", listing support in Gemini 2.5 Flash Live Preview: "we use non-blocking tools" is a statement about a model, not a vendor.

Two designs Ferrowick could defend​

Two architectures, with costs in different currencies.

Design A — hold before the irreversible commit point​

Move the commit point later, so ordinary cancellation is exact. No model-decided turn commit reaches a submission: Kelter restates the exact normalised command, takes an explicit confirmation, and the command enters HELD for a fictional twenty seconds, during which cancellation is exact because no external party has been told anything.

Three details carry the weight. The hold timer lives in the gateway, not the conversation, so a dropped connection neither releases nor cancels a held payment; a conversation-owned timer would turn every network event into a payment decision. The confirmation is — it names that exact normalised command, and any barge-in touching the confirming turn invalidates it. And it binds to the authenticated session rather than to voice recognition; the strongest current public statement on that is NIST's digital identity guidance, which says in section 3.2.3 that biometric comparison based on voice shall not be used.

A cheaper variant deserves naming: hold dispatch a few hundred milliseconds after the turn commits, and cancel silently if speech resumes. It costs no extra turn and no pending state, and on this timeline any window past +2.720 would have caught the correction. It is worth doing and it is not Design A — it moves the boundary slightly later without making cancellation exact, does nothing about narration, and the window a human self-correction needs is nearer two seconds than one, by which point you have paid Design A's latency without buying its certainty.

Costs. An extra turn on every release; a fixed delay before every submission; a new operator-visible pending state to explain and monitor; lower fluency, which is what operators complained about in the first place; and releases near the cut-off can miss the day.

Residual risk. The hold relocates the irreversible boundary — a cancellation one millisecond after expiry is back to a refusable recall — and it teaches operators that cancellation always works, so the rare post-hold case finds them least prepared. Confirming every release also trains people to confirm. On its own it fixes nothing about narration.

Ruling out voice biometrics creates a second residual risk rather than closing one. If a voice cannot identify a speaker, a spoken confirmation binds to an authenticated session on a device and nothing more — and the device is a hands-free handset, always listening, in a bay with a reversing lorry and other people in it. Anyone within earshot can say the confirming phrase, and Design A will treat it as Priya. The control rests on physical custody of an authenticated device: a much weaker claim than "confirmed by the operator", and the honest one. Buying attribution back costs the hands-free property the product exists for.

Design B — effect-authoritative narration​

Let the submission go, and never let any plane narrate a state it does not own. Kelter may describe its own domains freely — it stopped speaking, it stopped the model, it asked the tool to stop — and may describe the effect only by quoting the ledger's current state for that commandId. Late receipts are appended unconditionally; conversation delivery is a separate decision with a separate owner.

The first version of this was not finished, and the gap is hard to see because every individual rule is correct. At the barge-in the ledger says CANCEL_RACE_PENDING, so the honest sentence is "I asked, and I do not know yet" — exactly what the design wants. Priya hears a truthful uncertainty and goes back to work. At +4.100 the receipt arrives and is appended, as required. Whether it is ever said out loud falls to conversation scheduling, one of whose legal values is drop. Nothing obliged anyone to bring the answer back to the person told there wasn't one, and a design that replaces a confident falsehood with a permanent silence has improved the record, not the outcome.

So Design B carries one more rule. Any command in CANCEL_RACE_PENDING or SUBMITTED_UNCERTAIN about which an operator has been told an uncertain answer opens a commitment to report the resolution — a record outliving the turn, the connection, and the session, closing only when a resolution has been delivered to a person. It has an owner (payment operations), a fictional fifteen-minute bound, and a delivery path that must reach a delivery bay: a spoken follow-up on Kelter's next turn, a handset alert, or a phone call. A banner in the web application does not count, because Kelter exists precisely because Priya is not looking at one.

That also fixes the divergence detector, which was watching the wrong thing. Comparing turn state against ledger state catches the old defect; under Design B turn state says "cancel requested" while the ledger says "accepted" — two records that do not contradict each other, so the check may never fire. The trigger that does fire is a ledger state advancing after an unresolved spoken uncertainty, with no resolution delivered inside the bound.

Costs. An authoritative read on the critical path, adding latency and an availability dependency: if the ledger is unreachable Kelter must refuse to characterise effect state at all — and because the gateway records a command before submitting it, a ledger outage stops releases as well as narration, which near the cut-off is a supplier-relationship failure in the same currency Design A pays in. Kelter must sometimes say "I do not know yet", which measures badly in user testing. And every open uncertainty is now something a human must close, so Design B grows Ruth's queue twice over.

Residual risk. The money is gone. What Design B guarantees is narrower than a true account of the world: that no plane speaks about a domain it does not own, and that what it says about the effect is what Ferrowick has observed, while the ledger is readable. The detector remains a dependency — but its silent failure does not restore the original incident, because the spoken account still comes from the ledger and nobody is told a falsehood. It returns Priya to an uncertainty she is never brought back from: a different failure, a smaller one, still not acceptable.

The decision, and the argument still open​

Three things shipped, in the order they apply. Manual turn commit is on by default for the payment-release path, removing the first cause at the cost of the hands-free property. Design B applies to every release. Design A applies to a declared risk class — first-time payees, and amounts above a threshold Ruth Vance owns and reviews on a stated date rather than a constant somebody once typed.

Read that against Wednesday morning and it says something uncomfortable. Marlstone Aggregates was not a first-time payee; it was on a run approved on Monday, against a delivery Priya had just receipted. This release would have fallen outside the risk class, been handled by Design B alone, and the payment would still have gone out. Ferrowick accepted that deliberately, because holding every release was the price Tomas wanted and Dela refused. What the shipped answer buys is not impossibility: it is that next time Priya is told, inside fifteen minutes, by someone whose job it is to tell her — and that Kelter never says a sentence about money it has no standing to say.

Tomas still thinks A should be the default: no model-decided turn commit should ever reach an irreversible boundary. Dela still thinks A is a comfortable lie: every hold window has an outside, a design exact 99% of the time trains people to trust it in the 1%, and the durable fix is that no plane narrates state it does not own. They are both right about something, they still disagree, and the disagreement is written down with an owner and a review date. That is what a decision looks like when there is no correct answer.

The tests that separated a stopped conversation from a stopped effect​

Ferrowick's strongest visible success signal had been "the release completed and the operator did not complain". The incident satisfied it.

The replacement harness grades eight questions independently, per trial, across repeated trials. All results are fictional.

  1. Turn-commit correctness — the false-endpoint rate on utterances that contain a self-correction, the shape most test scripts omit.
  2. Barge-in to audible silence, measured to the client's reported playout cursor, not the server's cancellation acknowledgement.
  3. Prohibited narration — did Kelter ever describe an effect state it had not read from the ledger? Never offset by task completion.
  4. Unresolved uncertainty — was every spoken "I do not know yet" closed with a delivered resolution inside the time bound?
  5. Effect outcome — was an unintended payment submitted, accepted, recalled, or refused, and were duplicates created?
  6. Reconnection consistency — after resumption, does history agree with the playout cursor?
  7. Accessible completion — can the task be finished with push-to-talk, manual turn commit, and the typed path, at equal authority?
  8. Evidence reconstruction — can connection epoch, session, turn, response, command, receipt, and playout cursor be joined on aligned clocks without retaining raw audio by default?

Graders 1, 2, 6 and 7 need recorded speech and consented trials, so they live in the harness; 3, 4, 5 and 8 run against production records and need no audio. The harness supplies the cases nobody wants: packet loss and jitter, an LTE handover, a forced connection cap, a planned provider disconnection, a late receipt arriving after resumption — and the one the team had never tested, a cancellation that wins the race, so the correct outcome is produced through the same uncertain path.

Media telemetry joins the eighth grader through the standard browser statistics, and for a voice product the useful members are the audio ones: roundTripTime, packetsLost, jitter, jitterBufferDelay, plus concealedSamples, concealmentEvents, insertedSamplesForDeceleration and removedSamplesForAcceleration — the last four defined by the W3C WebRTC statistics API for audio and prohibited for video, which is what makes them the right instruments here. With those, "the model was slow" and "the network stretched the playout" stop being the same-looking incident. Raw audio retention is off by default; the causal chain lives in identifiers and timings.

Turn-taking is an accessibility decision​

Ferrowick's endpointing dial did not cause this incident — a second and a half of silence commits under any setting. But the dial has a separate cost worth naming on its own terms, because it decides who can use the product at all.

A threshold short enough to feel responsive is short enough to cut off someone who pauses to think, someone speaking a second language, someone using assistive technology, and someone with a speech disability or a stammer. Google's guidance about short thresholds fragmenting an utterance is a performance note; this story is what a fragmented utterance costs when the fragment is a command. Priya was not one of the people the dial disadvantages, and automatic turn detection still split her sentence. For the people it does disadvantage, that is not a rare event.

Manual turn commit is the mitigation, and it is an accessibility control before it is a safety one: a person who needs longer to speak should not have to win a race against a timer. The typed path in the site app carries the same authority rather than a degraded one; an equivalent path that cannot do the same thing is not an equivalent path.

Two constraints sat in the design from the start. Capture and recording consent are separate from microphone permission, and Kelter discloses that the counterpart is an AI system: the European Commission published guidelines on the AI Act's Article 50 transparency obligations on 20 July 2026, applying from 2 August 2026 and covering direct interaction with an AI system among other areas. That is a dated design constraint on a fictional product and emphatically not legal advice. And the handsets reach the media edge with a short-lived, backend-minted session credential, bound differently by each provider — either way it admits a device to media and authorises no payment. Every release is authorised at effect time against the authenticated session and the approved run, not against possession of a media token and not against a voice.

Transfer: the bid that closed the gate​

Ruth Vance tests whether a team learned a voice trick or a control model by moving the problem somewhere nobody can argue with the deadline.

A fictional energy-services company runs a voice assistant for plant energy managers, who offer blocks of flexible load into a wholesale market ahead of the operator's published gate closure. After closure the offer is binding: deliver the flexibility, or pay an imbalance charge. There is no equivalent of a recall request, because nobody has an incentive to grant one.

A manager says: "offer four megawatts for the evening block — no, wait, line two is back up tomorrow." The endpointer commits at the pause. The offer goes out as a non-blocking tool call. The barge-in cancels playout, generation, and the turn, and asks the market gateway to withdraw. Gate closure passed 900 ms ago.

Four questions, and a recap answers none of them:

  1. Name the domains. List every cancellation domain in that system, name its owner, and say what each acknowledgement proves and what it does not.
  2. Locate the commit point. Where exactly is the irreversible commit point, who owns it, and who may refuse a withdrawal?
  3. Design the control. Choose hold-before-commit or record-authoritative narration as the default; name the risk class where you would apply the other, and the cost you accept in the units you will be judged on — revenue, latency, human queue depth, or exposure.
  4. Falsify. Write the sentence your design assumes is true — "our withdrawal path stops the bid" — and describe the test that would show it false. If no test can show it false, you have not made an engineering claim.

"Add a confirmation step" is partial: it moves the commit point and says nothing about narration. "Make the timeout longer" is not an answer; every window has an outside. "Detect and alert" is half an answer, and the half it supplies is a dependency.

The surface changed from supplier payments to a capacity market. The rule did not:

A cancellation is a request. Only the plane that owns an effect may say what happened to it.

That is the same epistemic move as an earlier story in this collection, about a failover agent that mistook no answer for no action. Different mechanism, same error: treating the absence of a signal as the presence of a fact.

Where cancellation authority lives​

The incident fanned one interruption out to five owners. The topology below shows where responsibility for cancellation and for narration sits in each design, and marks the path from turn state to a spoken claim about money as blocked.

Loading interactive production topology…

Each enclosure carries its owner and its ordered steps, so the short version is what changed. The handset now reports the playout cursor as well as clearing it. The session owner truncates the assistant's turn to that cursor, and owns what was said rather than what happened. The orchestrator records that a cancellation was requested, and may describe only its own domains. The gateway owns the command lifecycle, the duplicate window, Design A's hold timer, the ledger, and unconditional receipt ingestion. Norwell Pay and the scheme own acceptance and whether a recall is honoured. Payment operations own the recall request, the open reporting commitments, and what a customer is told.

Ferrowick can now stop a conversation and say honestly what that did and did not stop. Kelter still cannot un-send a payment. Nothing can. What changed is that nobody is told otherwise, and nobody is left waiting for an answer that never comes.

Concepts in this story5 concepts
Cancellation domainRealtime multimodal · Interruption and cancellation

One region of state that a single cancellation actually reaches, together with the system that owns it. A live interaction has several — audio playout, model generation, turn and command intent, the tool request, and the external effect — and each acknowledges separately, at a different speed, with a different meaning. How fast an acknowledgement arrives tells you nothing about what it proves, and the domain that owns the effect may never acknowledge at all.

Committed turnRealtime multimodal · Turn-taking

The point at which a system decides a person has finished speaking and treats their words as an input it will act on. Silence thresholds and semantic endpointers make that decision probabilistic, so a committed turn is a guess about a human thought, not a fact about it — and an irreversible effect must not depend on one.

Unheard-output cursorRealtime multimodal · Session continuity

The boundary between output a system generated or sent and output a person actually heard. Realtime models emit audio faster than realtime, so an interruption leaves buffered speech that was delivered and never played. Conversation history must be reconciled to the played position, or a resumed session will confidently refer to a confirmation nobody received.

Irreversible commit pointDistributed systems · Effect safety

The moment in an effect’s lifecycle after which cancellation no longer exists, and the only remaining option is a separate compensating request that another party may refuse. Before it, cancellation can be exact because no external system has been told anything; after it, a cancellation is a negotiation. Holding an effect moves this point later; it never removes it.

Stale-result suppressionRealtime multimodal · Conversational tools

Dropping a late asynchronous result from the current conversation because its turn, topic, or command has moved on. It is the right conversational policy and the wrong records policy: a receipt must always enter the authoritative effect log, while whether it is spoken aloud is a separate decision with a separate owner.


Evidence and fiction note​

Fictional. Ferrowick, Kelter, Norwell Pay, Marlstone Aggregates, the Halden Row project, Priya Raghavan, Tomas Idrissi, Dela Okonkwo, Ruth Vance, the command lifecycle and its state names, the twenty-second hold, the fifteen-minute reporting bound, the duplicate window, the scheme cut-off, every timestamp, offset and latency, every amount and threshold, and every evaluation and harness result are invented for this story. They are internally consistent; they are not measurements, and they describe no real company, person, product, or payment scheme.

What the real sources support. All pages checked 2026-08-10.

  • OpenAI Realtime — the two turn-detection modes and the eagerness dial (VAD guide); server-side cancellation with response.cancelled, the instruction that the client note how much audio was played and send conversation.item.truncate with that position, the imprecision of transcript-to-audio alignment, and the 60-minute maximum session duration (conversation guide); server-side client secrets (WebRTC guide).
  • Google Gemini Live API — discarded generation on interruption, history retained to what was sent, and the activity-detection knobs (capabilities guide, 2026-08-05); the 15-minute audio session that applies "without compression", the unconditional roughly 10-minute connection, context-window compression extending sessions, GoAway, and two-hour resumption handles (session management, 2026-06-01); NON_BLOCKING behaviour, the three scheduling modes, manual tool-response handling, and the model-version limit on asynchronous calling (tool guide, 2026-06-01); single-use media credentials (ephemeral tokens, 2026-07-30).
  • Protocols and standards — cancellation as a late-arriving, ignorable request, and a disconnect as a cancellation signal over Streamable HTTP (MCP revision 2026-07-28); attempted-not-guaranteed remote cancellation (A2A 1.0.0); pending-request versus established-session teardown (RFC 3261); the voice-biometric prohibition (NIST SP 800-63B §3.2.3); the audio-only telemetry members and the prohibition of freeze and pause counters for audio (W3C WebRTC Statistics API); and the 20 July 2026 publication and 2 August 2026 application dates for the Article 50 obligations (European Commission guidelines entry, with covered areas on its policy page).

What the real sources do not support. They do not describe Ferrowick, its incident, its architecture, or its numbers, and they define no cancellation of a business effect a realtime tool call has already started — no source reviewed for this story does, and that absence is the story's central point rather than a gap in the research. They establish no payment scheme's rules, timings, cut-offs, or recall behaviour, and this story never claims a recall usually succeeds. Nothing here is legal advice.

Where providers differ. Both differences the story leans on are documented, and neither is a matter of degree. Reconciling history after an interruption is placed on the client by OpenAI, through a truncation event whose alignment it calls imprecise, and on the server by Google, which reconciles to what it already sent and asks the client to clear playback. And both cap a conversation without capping the same thing, or the same way: a 60-minute maximum session on one, versus two stacked ceilings on the other — a 15-minute audio session that context-window compression removes, inside a roughly ten-minute connection that it does not — so the durable limit is the one you cannot configure away, and reconnection is exceptional on one platform and routine on the other. Asynchronous tool scheduling is documented by Google and, on the realtime guides reviewed, not by OpenAI, and Google notes it exists on some Live models and not others. Any design assuming one provider's semantics apply to the other is assuming something no documentation supports.

What this story does not claim. It offers no medical, legal, or emergency-response guidance, depicts no safety-of-life system, and treats nothing as solved. The consequence here is money, materials, and wasted working days: a pour slipped and a groundworks crew stood idle. It shows no workaround for consent, recording, or disclosure obligations, because a design that routes around them is a worse design, not a cleverer one.