The rollback that left one deal in two model versions
MeridianClear is a fictional capital-markets operations company. Banks and investment firms use its platform to turn signed trading agreements and later amendments into structured terms that booking, risk, and client-operations systems can use.
Five finance terms are enough for this story. An agreement contains the original legal terms. An amendment changes them later. Structured terms are the machine-readable proposal extracted from those documents. Booking commits an approved operational revision. A collateral requirement is one downstream calculation that uses the booked terms.
MeridianClear built TermWeaver, an AI application that proposes structured terms from signed documents. It reduced repetitive reading, but it did not decide what a contract meant and it could not book anything.
Mira Sen (the AI platform architect responsible for immutable behavior manifests, inference-attempt evidence, routing, and recovery coordination) owned the path from document input to a reviewable proposal. A contract-operations analyst checked the proposal against the source. A second approver provided four-eyes authorization before booking. Legal Product owned clause-precedence policy. The Booking Platform owned committed revisions. Risk and Margin owned collateral calculations and client notices derived from those revisions.
How one revision normally moves through MeridianClear
The signed documents were authoritative for legal language. TermWeaver's output was a proposal. Its deterministic compiler proved that declared fields matched schema, referenced clauses existed, and required values were present. It did not prove that a model had preserved every natural-language precedence relationship.
Two people authorized booking, but their approval was not independent truth: both could be influenced by the same plausible proposal. Once committed, the booking ledger became operationally authoritative for downstream calculations. It did not replace the signed source as legal authority.
That distinction mattered because MeridianClear was moving from behavioral release C18 to C19.
Concepts in this story5 concepts
The complete compatibility surface that jointly determines a generative attempt’s behavior: model, tokenizer, template, adapter, context and reasoning policy, decoding, validator, runtime, precision, topology, provider, route policy, and evaluation release where applicable. Model weights alone do not identify it.
An application-owned immutable identifier that resolves to one behavior-bundle manifest before an attempt runs. Friendly aliases and requested or returned provider model names are useful evidence, but none is a complete or stable release identity by itself.
The supported period in which more than one behavior release or release-produced representation remains active. Equal schemas or locally valid releases do not prove that outputs produced under different semantic contracts can safely be combined.
The parentage connecting a request and each retry, fallback, or continuation to its exact behavior release, input artifacts, accepted output, authority receipts, and durable descendants. Human edits may add ancestry; they do not erase the earlier influence.
Evidence that every declared recovery plane has reached its intended boundary: future routing, in-flight work, cached or generated representations, committed application state, external artifacts, evaluation rows, and learning data. Restoring future traffic is one plane, not proof that rollback is complete.
The upgrade was larger than a model name
TermWeaver's release dashboard showed a friendly alias: contract-reasoner-current. The alias concealed the thing Mira actually operated.
A included the model snapshot, tokenizer, chat template, adapter, adaptation corpus, input assembly, context and truncation policy, reasoning policy, decoding settings, constrained grammar, validator, runtime, precision, serving topology, admission rules, provider and region, and the evaluation release that qualified the combination.
Changing any of those could change an answer. C18 and C19 therefore received immutable that bound the whole bundle, not just a weights file.
C19 was a legitimate upgrade. Its source-reading tests were stronger. It represented clause replacement and deletion as typed relationships instead of flattening every document revision into the latest visible value. On complete signed source, both releases produced the right answer for the locked full-source control.
The migration risk lived in a representation that was never described as an interface.
The useful summary C18 left behind
To avoid reprocessing hundreds of pages on every amendment, C18 stored a compact deal summary after an approved revision. The summary was useful within C18's assumptions. It was smaller, faster to load, and cheaper to send through the inference path.
Fixture SRC-4821-R5 contained three facts:
- the 2018 agreement set an amount at GBP 12 million;
- a 2024 amendment replaced it with GBP 8 million; and
- a 2026 amendment deleted that replacement, restoring the original GBP 12 million from 15 July 2026.
Before the final amendment arrived, C18 had produced summary SUM-771 under schema S6:
{
"schema": "term-summary/S6",
"deal_id": "D-4821",
"portfolio": "A",
"currency": "GBP",
"independent_amount": 8000000,
"effective_from": "2024-02-01",
"source_set_digest": "sha256:fictional-src-r4"
}
The row was valid. It represented the deal as C18 needed it at revision R4. But S6 retained no edge saying “GBP 8 million overrides a GBP 12 million base.” It preserved the current value and discarded how that value came to be current.
C19 expected its own prior representation to preserve that typed provenance. With full source, it did not need to trust the summary at all.
Three trials isolated the failure
At 09:41 on 18 July, the unsafe request path supplied SUM-771 and only the new amendment to C19. C19 saw a current amount of GBP 8 million and an amendment deleting a replacement relation it could not reconstruct. It proposed GBP 8 million.
The compiler accepted the proposal. The field was an integer, the currency was allowed, the source references existed, and schema S6 was complete. The compiler had proved exactly what it was designed to prove.
Mira refused to call this a bad-model incident until the controls were run together:
| Trial | Input and consumer | Output | Compiler | Semantic result |
|---|---|---|---|---|
| Pure C18 | Full SRC-4821-R5 → C18 | GBP 12m | Accept | Pass |
| Pure C19 | Full SRC-4821-R5 → C19 | GBP 12m | Accept | Pass |
| Mixed | C18 SUM-771 + 2026 amendment → C19 | GBP 8m | Accept | Fail |
If either pure trial failed, the team had a release defect. If the compiler rejected the mixed output, they had a contained incompatibility. If SUM-771 retained the base and typed override edge, the proposed mechanism was false.
None of those happened. The failure existed only in : work from distinct behavioral releases was simultaneously live, and an intermediate representation produced under one release crossed into another without an explicit compatibility contract.
The new model did not “forget” the agreement. The old model was not generally wrong. C19 received a representation that was valid for C18 and insufficient for C19. Local validity did not compose into system validity.
The one deal was already more than one row
The proposal became TERM-990. A contract-operations analyst compared the changed-looking section with the amendment. A second approver authorized it. Neither reconstructed the deleted relationship from the 2018 and 2024 documents because the screen emphasized the plausible GBP 8 million proposal.
The Booking Platform committed a new revision. Risk and Margin recalculated exposure. A collateral workflow prepared a client request. The approved row entered an evaluation sample, and the sample joined an adaptation-data snapshot.
This was the complete :
SRC-4821-R4
→ C18 attempt
→ SUM-771
SRC-4821-R5 + SUM-771
→ C19 attempt
→ TERM-990
→ compiler receipt
→ analyst approval + second approval
→ booked revision
→ exposure and collateral views
→ client notice
→ evaluation row
→ adaptation snapshot
That graph was evidentiary only because each attempt had an immutable receipt. At minimum it bound the attempt and parent IDs, deal and revision, complete behavior-manifest digest, requested and returned provider model IDs, source-snapshot digest, route-policy digest, parent artifact IDs and content digests, output artifact and digest, validation receipt, and accepted term-set ID. If a reviewer edited a proposal, MeridianClear added an editor receipt and retained the generated artifact as an ancestor. Matching strings in the final prose was explicitly rejected as lineage.
An inference attempt was no longer disposable telemetry once its output influenced an authoritative or external system. Its descendants needed identities and state transitions of their own.
The rollback was correct—and incomplete
At 12:08 UTC, Mira changed future routing back to C18. That was the right first action. It prevented new C19 attempts while the team investigated.
It did not make C19's completed work disappear.
The recovery ledger recorded each population separately:
| Population at rollback | Required action |
|---|---|
| 17 active C19 attempts | Cancel 11 before output; quarantine 6 completed but uncommitted proposals |
| 84 C19 attempts that consumed a C18 summary | Resolve producer manifest, consumer manifest, source revision, and decision lineage |
| 23 attempts past compiler and approval | Re-derive from full source and run dual review; identify 5 divergent results |
| 5 divergent booked revisions | Forward-correct while preserving the inactive history |
| 3 external collateral requests | Reconcile with operations; recompute all 5 affected exposure views |
| 2 client notices already sent | Issue superseding notices linked to the originals |
| Canary evidence | All 23 accepted mixed attempts polluted the canary success aggregate; rebuild by immutable manifest and representation epoch, with mixed attempts reported separately |
| 23-row adaptation snapshot | Revoke it and prove no descendant release was promoted from it |
The team also attached a producer manifest and semantic contract to every retained summary. A broad cache clear would have removed some convenient copies. It would not have identified booked descendants, external requests, exported evaluation rows, or notices a client had already read.
This is : not merely sending future traffic to an older route, but driving every affected attempt and descendant toward an explained, repaired, quarantined, revoked, or explicitly residual state.
The distinction prevented a false victory. If an external extract could not be traced, MeridianClear recorded residual risk. It did not claim total eradication.
The guardrail that should have been boring
The unsafe handler selected a target release and reused whatever prior summary was convenient. Mira replaced that assumption with an explicit coexistence decision:
async function prepareRevision(input: DealRevisionRequest) {
const [prior, target, source] = await Promise.all([
summaries.findApproved(input.dealId),
releases.resolveImmutable(input.behaviorReleaseId),
legalSourceAuthority.resolveAndLockCompleteSet({
dealId: input.dealId,
asOfRevision: input.asOfRevision,
}),
]);
const attempt = await attempts.open({
dealId: input.dealId,
behaviorReleaseId: target.id,
sourceSnapshotId: source.snapshotId,
sourceSnapshotDigest: source.digest,
});
if (prior.behaviorReleaseId === target.id) {
return interpretAmendment({amendment: input.amendment, prior, source, target, attemptId: attempt.id});
}
return coexistencePolicy.choose({
rebuild: () => interpretFullSource({source, target, attemptId: attempt.id}),
migrate: () => openDualRunMigration({input, prior, source, target, attemptId: attempt.id}),
descendants: await lineage.findByDeal(input.dealId),
});
}
The code resolved and locked the complete document set from the legal source authority by deal and as-of revision; it did not trust a caller-supplied source snapshot. The attempt receipt bound that snapshot's immutable ID and digest before either interpretation path ran.
It could then detect a release boundary and force one of two controlled paths. It could not decide the legal meaning of a clause, migrate descendants by itself, guarantee that a nondeterministic rerun would be identical, retain a retiring provider indefinitely, or decide how much untraceable residual risk the business should accept.
Those remained accountable decisions.
Design A: re-derive from full source
Under the first correction, every authoritative revision assembled a complete immutable source snapshot and one complete behavior manifest. TermWeaver generated from that snapshot, the compiler checked declared structure, reviewers compared the proposal with signed source, and booking recorded the attempt lineage. Canary descendants stayed quarantined until the revision cleared review.
This bought rebuildability, faster retirement of an old bundle, and a smaller legacy fleet. A cross-version summary could no longer become an invisible dependency.
It cost input tokens, latency, document-access capacity, privacy exposure, and reviewer time. It could delay irreversible descendants. It also left real residual risk: an incomplete source snapshot, a shared compiler blind spot, model variation, or a semantic question requiring legal adjudication. A correction after an external effect still had to move forward; re-derivation did not reverse time.
Design B: pin the deal, then migrate it explicitly
Under the second correction, ordinary revisions remained pinned to one long-lived behavioral epoch. Migration locked the complete source, ran old and new bundles independently, produced a typed semantic diff, sent disagreement to accountable adjudication, and changed the deal epoch only with a descendant plan.
This bought semantic continuity across ordinary amendments and made migration a named operational event.
It cost long-lived bundles, runtimes, provider contracts, validators, evaluation suites, backports, dual-run capacity, human review, and retirement machinery. It left different residual risk: an inherited legacy defect, a blind spot shared by both versions, a provider forcing retirement before migration completed, or a bad adjudication decision.
Neither design was “the safe one.” The decision depended on source size, amendment frequency, provider retention, privacy limits, acceptable latency, and the cost of maintaining old behavior. What mattered was that coexistence stopped being accidental.
What the release council changed
MeridianClear stopped asking, “Which model served this request?” The release review required a stronger set of evidence:
- the immutable behavioral release identity for every producer and consumer;
- a semantic contract for every representation allowed to cross between them;
- pure-old, pure-new, and mixed-path control trials;
- attempt lineage through every authoritative, external, evaluation, and adaptation descendant;
- a rollback plan with convergence criteria, not only a routing command; and
- an explicit residual-risk statement for anything that could not be traced or reversed.
Model providers expose pinned or versioned identities, moving aliases, and lifecycle or deprecation schedules; some provider identifiers are dated. Serving systems support canary traffic splits, and adapter-serving runtimes can load multiple adaptations dynamically. Those capabilities make coexistence possible. They do not make intermediate representations compatible or descendants self-repairing.
Your turn: move the incident into retail
A retailer is migrating its product-catalog interpreter from C18 to C19. C18 produced a normalized supplier summary for one product. A supplier then sends a new amendment, and the catalog service combines that C18 summary with the amendment under C19. The resulting artifact feeds search ranking, fulfillment rules, and a future training snapshot.
Before choosing a correction, write a short architecture decision that answers all five questions:
- Controls: What pure-C18, pure-C19, and mixed-path trials would distinguish a generally broken release from an unsafe composition? State the expected output and a falsifier for each.
- Authority: Which system is authoritative for supplier meaning, which system may propose normalized attributes, who may approve a catalog revision, and which platform commits it?
- Descendants: If future traffic returns to C18, which search indexes, fulfillment decisions, customer-visible pages, evaluation rows, and training artifacts remain unrepaired?
- Architecture: Would you re-derive each accepted revision from complete supplier source under one manifest, or pin products to an epoch and perform typed migration? Defend the choice for this retailer's amendment rate and source size.
- Cost and residual risk: Price the choice in latency, input tokens, infrastructure, reviewer work, legacy support, and retirement effort. Name at least one risk the design accepts rather than pretending to eliminate.
A Professional answer must preserve the same causal discipline as the MeridianClear investigation: both pure controls can pass while the mixed path fails; source, proposal, approval, commit, and descendant repair belong to different authorities; rollback of routing does not repair materialized descendants; and neither architecture removes every risk.
The sentence Mira carried to the next migration review was narrower and more useful:
A behavioral release is the whole path that can change an answer. If two releases coexist, version the representations between them. If one is rolled back, follow its effects until every descendant has an explained state.
Concepts in this story5 concepts
The complete compatibility surface that jointly determines a generative attempt’s behavior: model, tokenizer, template, adapter, context and reasoning policy, decoding, validator, runtime, precision, topology, provider, route policy, and evaluation release where applicable. Model weights alone do not identify it.
An application-owned immutable identifier that resolves to one behavior-bundle manifest before an attempt runs. Friendly aliases and requested or returned provider model names are useful evidence, but none is a complete or stable release identity by itself.
The supported period in which more than one behavior release or release-produced representation remains active. Equal schemas or locally valid releases do not prove that outputs produced under different semantic contracts can safely be combined.
The parentage connecting a request and each retry, fallback, or continuation to its exact behavior release, input artifacts, accepted output, authority receipts, and durable descendants. Human edits may add ancestry; they do not erase the earlier influence.
Evidence that every declared recovery plane has reached its intended boundary: future routing, in-flight work, cached or generated representations, committed application state, external artifacts, evaluation rows, and learning data. Restoring future traffic is one plane, not proof that rollback is complete.
MeridianClear, its people, systems, incident, fixtures, and metrics are fictional. Technical grounding: provider documentation on Claude model identifiers and aliases and Gemini model identities and lifecycles; KServe documentation on canary rollout and rollout observability; NVIDIA Dynamo documentation on dynamic LoRA adapter serving; and DriftBench research on infrastructure drift across LLM serving configurations. These sources establish relevant production mechanisms, not the fictional failure.