Skip to main content

The package no technician had inspected

· 14 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerPackProof inspected six camera views. The missing evidence was an independent physical inspection of the package.

MorrowVale Foods is a fictional refrigerated-food manufacturer. Its packaging lines print and apply the required label to each sealed retail package before cases move to the warehouse. A wrong artwork revision, unreadable allergen text, missing language panel, or damaged lot code can place that package on hold.

MorrowVale built PackProof, a fictional multimodal inspection service, to help its quality team find those problems. Six synchronized cameras photographed every visible side of one sealed package. PackProof read those images beside the package's approved artwork specification and proposed an anomaly score, evidence quality, suspected issue types, and a priority for physical inspection.

PackProof inspected every package's six camera views. What did not happen for most low-risk packages was the independent physical inspection that MorrowVale had named as its authoritative evidence.

Leena Rao, the quality systems engineer responsible for PackProof's inspection evidence and training pipeline, had kept the model's authority deliberately narrow. PackProof could propose where scarce technician attention should go. A deterministic inspection-policy service chose the route. Warehouse systems owned release and hold. Only a trained quality technician using MorrowVale's fictional QL-4 protocol could create a physical pass or fail observation.

That was the intended design. The failure appeared later, where evidence became training data.

The rollback that left one deal in two model versions

· 14 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about a correct rollback that did not repair the work already produced by two behavioral releases.

MeridianClear is a fictional capital-markets operations company. Banks and investment firms use its platform to turn signed trading agreements and later amendments into structured terms that booking, risk, and client-operations systems can use.

Five finance terms are enough for this story. An agreement contains the original legal terms. An amendment changes them later. Structured terms are the machine-readable proposal extracted from those documents. Booking commits an approved operational revision. A collateral requirement is one downstream calculation that uses the booked terms.

MeridianClear built TermWeaver, an AI application that proposes structured terms from signed documents. It reduced repetitive reading, but it did not decide what a contract meant and it could not book anything.

Mira Sen (the AI platform architect responsible for immutable behavior manifests, inference-attempt evidence, routing, and recovery coordination) owned the path from document input to a reviewable proposal. A contract-operations analyst checked the proposal against the source. A second approver provided four-eyes authorization before booking. Legal Product owned clause-precedence policy. The Booking Platform owned committed revisions. Risk and Margin owned collateral calculations and client notices derived from those revisions.

The green test that had never been told the rule

· 20 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the difference between a test that records output and a test that states a rule.

Kestrelway Transit is a fictional regional transit operator. Riders tap a bank card or a phone at a gate; there is no ticket to buy and no fare to choose in advance.

Kestrelway built FareBridge to price one rider's complete service day, which runs from 03:00 to 03:00. It takes everything a rider did in that window and produces one figure.

Ola Fadeyi was the fare policy analyst, and owned the published fare rules, their wording, and the date each took effect. Hugo Bertrand was the fare platform engineer who owned FareBridge, its regression suite, and the gate that decided whether a change could ship. Grace Mbeki led revenue operations and owned daily settlement reconciliation.

FareBridge decided what a day cost. Settlement decided that the charge was sent — once.

The healthy cell that no request could reach

· 23 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the difference between running a system and changing one.

Slatebridge is a fictional cloud dispatch platform for independent home-services contractors — the heating, plumbing, and electrical companies with four to sixty technicians that arrive at your door in a small van. Roughly 2,400 fictional contractor companies and 19,000 technicians use it.

Slatebridge Desk is the browser app where a dispatcher builds tomorrow's route. Slatebridge Field is the phone app where a technician opens that route at shift start and marks each job done on the customer's driveway. Underneath, Slatebridge runs five cells, cell-01 through cell-05, with one shared entry layer in front of all five: Gatehouse.

Beatriz Quintero was the staff platform engineer who owned Gatehouse and Tenant Control, the service that creates contractors and decides where they live. Ivo Krastev was the site reliability engineer on call the morning this story happens. Wes Halloran ran support, so he learned the size of a problem before the dashboards did.

Slatebridge promised contractors one thing above all: from 06:30 local time onward, a technician can open the day's route in one tap. A technician who cannot open a route does not start driving.

The interruption that stopped the sentence and not the payment

· 40 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about barge-in, separate cancellation domains, and why a spoken assistant must never describe an effect it does not own.

Ferrowick is a fictional construction-operations platform used by mid-sized building contractors. Customers run whole projects on it: purchase orders, delivery receipting, retention tracking, subcontractor valuations, and the weekly supplier payment run. Ferrowick holds no customer money; payments are initiated through Norwell Pay, a fictional licensed payment provider that submits into a bank scheme.

Kelter is Ferrowick's realtime site assistant, and it exists because of one physical fact. The person who knows what actually arrived is standing in a delivery bay in gloves and hearing protection, beside a reversing lorry, with a clipboard in one hand. She is not going to open a laptop. Kelter runs on a site handset or a paired headset, connects over WebRTC to a regional media edge, listens, and speaks.

A payment passes three decisions, made by different people at different times. On Monday the project manager approves the week's run in the web application, deciding what may be paid at all. During the week the site operations lead receipts deliveries, deciding what arrived. When retention conditions are met she releases the matching payment by voice, deciding what is sent now and in what order.

That third decision matters, because a project's available balance is finite and sequenced. Releasing early does not move money to the wrong week; it consumes the headroom a later payment depends on. A supplier who is not paid does not deliver, and a crew that arrives to an empty bay stands idle at full cost.

Priya Raghavan is Site Operations Lead on the fictional Halden Row project; she receipts deliveries and releases approved payments. Tomas Idrissi is a Staff Engineer who owns Ferrowick's realtime interaction plane — media session, connection epoch, turn state, playout, interruption handling. Dela Okonkwo is a Principal Engineer who owns the payment-effect plane — the gateway, the operation ledger, the Norwell Pay integration. Ruth Vance is Head of Payment Operations, and owns what a customer is told when the state of a payment is uncertain.

Tomas and Dela both built correct systems. The sentence that caused this incident lived in Tomas's plane and was about Dela's, and nobody had ever decided which of them owned it.

The replay that reproduced every event and changed the bill

· 24 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Data systems storyAdvancedA fictional retail-energy incident in which a replay reproduced every event in order and produced a number no consistent rebuild could produce.

Larkspur Energy is a fictional retail electricity supplier. It sells Larkspur Flex, a plan that prices consumption by time-of-use band and adds a demand charge derived from the largest half-hour a site draws inside a peak band. Its customers are homes and small businesses. It does not own meters: a fictional metering agent, Gridpost Metering, reads them and publishes the data. It does not own the market either — it submits a daily settlement position to a regional market operator and lives with what the operator settles.

Sanne Vermeer (streaming platform engineer) owns the two stream jobs that turn meter reads into money, and the tooling that replays them. Dov Aharoni (settlement analyst) owns the daily market submission and authorises backfills when metering data goes missing. Marta Oyelaran (billing operations lead) owns customer notices, corrections, and the policy that decides which plan a site is eligible for. Marta owns the consequence of this incident and was not part of the repair that caused it.

How a half-hour of electricity becomes a bill​

Loading the normal Larkspur business flow…

Larkspur Flex is defined by five rules. All five are Larkspur's own, or its market's; none describes a real market.

RuleValue
Rate classesOFF_PEAK, SHOULDER, PEAK, CRITICAL_PEAK
Band adjustment applied to each rated interval when computing the demand metric1.01 / 1.02 / 1.06 / 1.09 respectively
Coincident peakPer half-hour, sum the site's adjusted kWh across its meter points; take the maximum over half-hours classed PEAK or CRITICAL_PEAK; double it to express kW
Plan eligibilityA rolling twelve-month coincident peak above 25 kW moves the site off Flex onto a demand tariff, with 14 days notice
Market submissionTrading day T is submitted on T+1, and the revision window closes on the 15th of the following month

A site can have several meter points — a main supply, a dedicated circuit, a charger — and the coincident peak is a property of the site, computed across all of them. That detail is the whole story, and it looks like nothing.

The right model that arrived too late

· 12 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about correctly routed reasoning work that consumed its useful deadline.

IonWeave Semiconductor is a fictional chip manufacturer. Its fabrication plants run hundreds of tightly controlled tools that deposit, etch, measure, and clean material during chip production. A stopped tool is expensive. An unsafe restart is worse.

IonWeave built ToolSage, an internal AI assistant for line engineers. When a fabrication tool raised an alarm, ToolSage received a sealed machine-state snapshot, maintenance notes, and the approved runbook revision. It drafted a cited troubleshooting plan. It did not clear interlocks, change a recipe, or send commands to equipment. It had no equipment credentials.

Yuna Park (the inference reliability lead responsible for model routing, reasoning policy, and admission) owned the path from a ToolSage request to a completed model response.

A line engineer checked the response against the machine state shown by the tool. An equipment owner alone could authorize an approved runbook step or take manual control. The useful outcome was therefore not “model produced text.” It was “owner received enough verified information to decide safely while the incident was still actionable.”

How an alarm normally becomes a decision​

Loading the ToolSage journey…

The Tool Control Gateway sealed the alarm code, interlock bit, recent sensor state, maintenance notes, tool support status, recipe revision, and current runbook. ToolSage then applied route policy TS-31 to those authoritative fields.

Known routine alarms used Torch-Triage-12B/r2026.07.2, prompt template TRC-5, with no adaptive reasoning and at most 600 output tokens. Novel recipe alarms and active safety interlocks used Torch-Reason-70B/r2026.07.1, also with TRC-5, with an adaptive 800–8,000-token and at most 1,000 output tokens. Both ran on the fictional IonServe/r4.8.1 runtime.

TS-31 did not ask a model whether an interlock was active. It read interlock_active, supported-tool state, and recipe revision from the gateway and registry. Model utility came after hard eligibility. If a request was eligible, the selected model proposed a plan; the engineer verified it; the equipment owner decided what happened next.

That was the system Yuna took into a recipe ramp.

The valid zone that could not hold the pallet

· 12 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the boundary between a well-formed proposal and a warehouse commitment.

VelaFresh Logistics is a fictional cold-chain warehouse operator. Its receiving teams move inbound pallets from temperature-controlled trailers into storage zones before those pallets continue to customer orders.

VelaFresh built DockPilot to help with the first decision. A receiver scanned a pallet, added a short exception note, and received a proposed zone. DockPilot used one model call to balance ordinary operating preferences: avoid unnecessary handling, keep urgent pallets near outbound doors, and make sensible use of the zones Zone Control said were open.

Noor Bakshi was the warehouse automation engineer who owned DockPilot and its handoff to the warehouse-management system, or WMS. DockPilot could propose a location. The WMS was the component that reserved that location and released a forklift putaway task.

That division mattered even on a normal day.

The export that passed every permission check

· 26 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about composed agent authority, effect-bound approval, and releasing exact bytes to the right person.

Harborlight People is a fictional workforce-management SaaS company. Its employer customers use the platform for payroll, scheduling, employee relations, support, and access administration. They can also configure a workflow through which current and former workers request copies of their personal data.

The employer customer defines the request policy, the source categories in scope, and who may make the final release decision. Harborlight operates the software and a managed privacy-operations team under that customer-defined policy. The process in this story is Harborlight's fictional design, not a universal legal requirement.

Harborlight handled about 3,800 worker-data requests per month. Most looked simple from the request portal: prove who you are, describe the employment period, wait while the records are assembled, and collect a package from an authenticated portal.

The work behind that path was not a single database query.

A worker might have a legal name, a preferred name, a former surname, several email addresses, and more than one worker identifier after a rehire or contractor conversion. A support ticket might mention a person without being about that person. An attachment might be linked to a ticket whose requester and uploader were neither its subject. A fixed join could recover the obvious records. It could not reliably resolve every free-text reference, copied attachment, or derived record.

That was why Harborlight built Lumen, a bounded AI privacy worker. Lumen read case-scoped records through narrow adapters, resolved aliases and free-text references, proposed which records concerned the requester, explained the evidence behind each proposal, and suggested redactions. Lumen could produce candidates. It did not own a person's stable identity, record-subject truth, case policy, approval, delivery credential, or the release of bytes.

Mara Chen, the Senior Privacy Operations Specialist responsible for defining request scope and resolving ambiguous record matches, supervised the managed workflow. Under each employer's policy, she could approve an exact release.

The distinction sounded conservative enough: Lumen proposed; Harborlight decided.

The failover agent that mistook no answer for no action

· 18 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about durable AI workers, uncertain cloud operations, and recovering external truth.

Northstar Ledger (a fictional B2B commerce-infrastructure company) routes checkout and inventory updates for regional grocery chains. Its customers do not run one Northstar checkout page. They use Northstar behind their own mobile apps, self-checkout kiosks, and delivery sites.

At the Saturday peak, those systems send roughly 48,000 order operations per minute. A short outage creates queues. An incorrect recovery can create something worse: orders accepted through a path that the company no longer believes is authoritative.

Northstar stored checkout state in an Amazon Aurora PostgreSQL global database. The primary Region served writes. A secondary Region stayed ready for disaster recovery. The operating rule was strict: Northstar could advertise only the writer confirmed by Aurora's current topology.

Imani (the Staff Reliability Engineer responsible for Northstar's database-recovery control plane) had spent the previous six months reducing the time between an incident page and a safe recovery decision.

Her team built Relay, a bounded AI incident worker. Relay read approved telemetry, compared an incident with reviewed runbooks, and proposed a typed recovery plan. It could not promote a database or change traffic. A deterministic policy service checked every proposal. A human incident commander then approved the exact target and the maximum tolerated data loss before any effect was allowed.

This mattered because Relay was useful precisely where incidents were messy. It could gather replication lag, recent deployments, health probes, and runbook constraints in seconds. It could explain why one recovery target fitted the evidence better than another.

It could not make an ambiguous cloud operation unambiguous by thinking harder.

The fine-tune that answered with last week’s policy

· 10 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about keeping an AI answer tied to current policy.

Ashvale Benefits (a fictional employee-benefits administration company) helps mid-sized employers manage health-plan enrollment, eligibility rules, and employee policy questions. It had eleven days before its busiest open-enrollment season.

The company had built an AI support assistant inside its member portal. The assistant was supposed to answer routine questions from each employer’s approved policy handbook, show the supporting section, and hand uncertain cases to the support team.

During open enrollment, that team expected eighty thousand questions. Most would be some version of the same thing: Does my plan cover this? When does coverage begin? Which form do I need?

The business goal was simple. Answer the routine questions immediately. Send the hard ones to a human. Never invent a benefit that did not exist.

There was one more rule from legal: every answer had to point to the policy section that supported it.

The retry that made the outage worse

· 11 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about retry ownership across an AI risk-decision path.

Kiteframe Pay (a fictional payment-risk infrastructure company) helps online marketplaces decide whether a card order should be approved, declined, or reviewed by a human. Merchants call Kiteframe during checkout, before they capture money or promise inventory to a shopper.

Most orders never need generative AI. Deterministic rules settle obvious cases in under 80 milliseconds. The difficult eight percent—new devices, unusual delivery patterns, sparse account histories—enter Aster, Kiteframe's AI-assisted risk analyst.

Aster was not a chatbot and could not approve a payment. It gathered approved evidence, used a language model to produce a typed risk recommendation with cited signals, and passed that recommendation to a deterministic merchant-policy engine. Without Aster, those ambiguous orders went to a manual-review queue. During a large sale, that queue could grow faster than Kiteframe's analysts could empty it.

At 02:13 on Tuesday, model latency rose sharply. Seven minutes later, Kiteframe's entire checkout success rate had fallen from 99.4% to 71%.

The model provider was recovering.

Kiteframe's retries would not let it.

The shadow model judged by the old model’s evidence

· 17 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about a shadow model evaluated on field evidence selected by the incumbent.

BrambleGrid is a fictional grid-asset inspection company. Regional electric utilities hire it to combine drone imagery, asset records, and certified field inspection so scarce crews visit pole-top equipment that most needs a closer look.

A drone could photograph thousands of connector assemblies in a morning. A physical visit was different. It needed access coordination, a qualified crew, and enough time to perform the same close-range inspection protocol on every asset. BrambleGrid had capacity for 240 flexible visits per week. Twenty always remained reserved for manual safety reports that did not come from a model.

That made field inspection both an operating resource and an evidence resource.

BrambleGrid's ML-assisted product was called Spanwatch. For one recent connector evidence snapshot, Spanwatch estimated the probability that a protocol-C3 inspection within seven days would find an actionable connector condition. It did not predict an outage, and it could not authorize a repair.

Dalia Moravec (the ML reliability engineer responsible for calibration, evaluation cohorts, and launch recommendations) owned Spanwatch's model evidence.

Jon Ibarra (the field-inspection planning lead responsible for the weekly 240-visit capacity ledger) owned which proposed visits entered the field workflow.

Mara Venn (the lead asset-integrity engineer responsible for protocol-C3 adjudication and label revision) owned the final condition record. A utility duty engineer—not Spanwatch, Dalia, Jon, or Mara—separately decided whether to restrict or repair an asset.

The boundaries looked fussy until the week BrambleGrid tried to replace its model.