Skip to main content

7 posts tagged with "LLM engineering"

Building, adapting, evaluating, and operating language-model applications.

View All Tags

The rollback that left one deal in two model versions

· 14 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about a correct rollback that did not repair the work already produced by two behavioral releases.

MeridianClear is a fictional capital-markets operations company. Banks and investment firms use its platform to turn signed trading agreements and later amendments into structured terms that booking, risk, and client-operations systems can use.

Five finance terms are enough for this story. An agreement contains the original legal terms. An amendment changes them later. Structured terms are the machine-readable proposal extracted from those documents. Booking commits an approved operational revision. A collateral requirement is one downstream calculation that uses the booked terms.

MeridianClear built TermWeaver, an AI application that proposes structured terms from signed documents. It reduced repetitive reading, but it did not decide what a contract meant and it could not book anything.

Mira Sen (the AI platform architect responsible for immutable behavior manifests, inference-attempt evidence, routing, and recovery coordination) owned the path from document input to a reviewable proposal. A contract-operations analyst checked the proposal against the source. A second approver provided four-eyes authorization before booking. Legal Product owned clause-precedence policy. The Booking Platform owned committed revisions. Risk and Margin owned collateral calculations and client notices derived from those revisions.

The right model that arrived too late

· 12 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about correctly routed reasoning work that consumed its useful deadline.

IonWeave Semiconductor is a fictional chip manufacturer. Its fabrication plants run hundreds of tightly controlled tools that deposit, etch, measure, and clean material during chip production. A stopped tool is expensive. An unsafe restart is worse.

IonWeave built ToolSage, an internal AI assistant for line engineers. When a fabrication tool raised an alarm, ToolSage received a sealed machine-state snapshot, maintenance notes, and the approved runbook revision. It drafted a cited troubleshooting plan. It did not clear interlocks, change a recipe, or send commands to equipment. It had no equipment credentials.

Yuna Park (the inference reliability lead responsible for model routing, reasoning policy, and admission) owned the path from a ToolSage request to a completed model response.

A line engineer checked the response against the machine state shown by the tool. An equipment owner alone could authorize an approved runbook step or take manual control. The useful outcome was therefore not “model produced text.” It was “owner received enough verified information to decide safely while the incident was still actionable.”

How an alarm normally becomes a decision​

Loading the ToolSage journey…

The Tool Control Gateway sealed the alarm code, interlock bit, recent sensor state, maintenance notes, tool support status, recipe revision, and current runbook. ToolSage then applied route policy TS-31 to those authoritative fields.

Known routine alarms used Torch-Triage-12B/r2026.07.2, prompt template TRC-5, with no adaptive reasoning and at most 600 output tokens. Novel recipe alarms and active safety interlocks used Torch-Reason-70B/r2026.07.1, also with TRC-5, with an adaptive 800–8,000-token and at most 1,000 output tokens. Both ran on the fictional IonServe/r4.8.1 runtime.

TS-31 did not ask a model whether an interlock was active. It read interlock_active, supported-tool state, and recipe revision from the gateway and registry. Model utility came after hard eligibility. If a request was eligible, the selected model proposed a plan; the engineer verified it; the equipment owner decided what happened next.

That was the system Yuna took into a recipe ramp.

The valid zone that could not hold the pallet

· 12 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about the boundary between a well-formed proposal and a warehouse commitment.

VelaFresh Logistics is a fictional cold-chain warehouse operator. Its receiving teams move inbound pallets from temperature-controlled trailers into storage zones before those pallets continue to customer orders.

VelaFresh built DockPilot to help with the first decision. A receiver scanned a pallet, added a short exception note, and received a proposed zone. DockPilot used one model call to balance ordinary operating preferences: avoid unnecessary handling, keep urgent pallets near outbound doors, and make sensible use of the zones Zone Control said were open.

Noor Bakshi was the warehouse automation engineer who owned DockPilot and its handoff to the warehouse-management system, or WMS. DockPilot could propose a location. The WMS was the component that reserved that location and released a forklift putaway task.

That division mattered even on a normal day.

The export that passed every permission check

· 26 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about composed agent authority, effect-bound approval, and releasing exact bytes to the right person.

Harborlight People is a fictional workforce-management SaaS company. Its employer customers use the platform for payroll, scheduling, employee relations, support, and access administration. They can also configure a workflow through which current and former workers request copies of their personal data.

The employer customer defines the request policy, the source categories in scope, and who may make the final release decision. Harborlight operates the software and a managed privacy-operations team under that customer-defined policy. The process in this story is Harborlight's fictional design, not a universal legal requirement.

Harborlight handled about 3,800 worker-data requests per month. Most looked simple from the request portal: prove who you are, describe the employment period, wait while the records are assembled, and collect a package from an authenticated portal.

The work behind that path was not a single database query.

A worker might have a legal name, a preferred name, a former surname, several email addresses, and more than one worker identifier after a rehire or contractor conversion. A support ticket might mention a person without being about that person. An attachment might be linked to a ticket whose requester and uploader were neither its subject. A fixed join could recover the obvious records. It could not reliably resolve every free-text reference, copied attachment, or derived record.

That was why Harborlight built Lumen, a bounded AI privacy worker. Lumen read case-scoped records through narrow adapters, resolved aliases and free-text references, proposed which records concerned the requester, explained the evidence behind each proposal, and suggested redactions. Lumen could produce candidates. It did not own a person's stable identity, record-subject truth, case policy, approval, delivery credential, or the release of bytes.

Mara Chen, the Senior Privacy Operations Specialist responsible for defining request scope and resolving ambiguous record matches, supervised the managed workflow. Under each employer's policy, she could approve an exact release.

The distinction sounded conservative enough: Lumen proposed; Harborlight decided.

The failover agent that mistook no answer for no action

· 18 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyProfessionalA fictional production incident about durable AI workers, uncertain cloud operations, and recovering external truth.

Northstar Ledger (a fictional B2B commerce-infrastructure company) routes checkout and inventory updates for regional grocery chains. Its customers do not run one Northstar checkout page. They use Northstar behind their own mobile apps, self-checkout kiosks, and delivery sites.

At the Saturday peak, those systems send roughly 48,000 order operations per minute. A short outage creates queues. An incorrect recovery can create something worse: orders accepted through a path that the company no longer believes is authoritative.

Northstar stored checkout state in an Amazon Aurora PostgreSQL global database. The primary Region served writes. A secondary Region stayed ready for disaster recovery. The operating rule was strict: Northstar could advertise only the writer confirmed by Aurora's current topology.

Imani (the Staff Reliability Engineer responsible for Northstar's database-recovery control plane) had spent the previous six months reducing the time between an incident page and a safe recovery decision.

Her team built Relay, a bounded AI incident worker. Relay read approved telemetry, compared an incident with reviewed runbooks, and proposed a typed recovery plan. It could not promote a database or change traffic. A deterministic policy service checked every proposal. A human incident commander then approved the exact target and the maximum tolerated data loss before any effect was allowed.

This mattered because Relay was useful precisely where incidents were messy. It could gather replication lag, recent deployments, health probes, and runbook constraints in seconds. It could explain why one recovery target fitted the evidence better than another.

It could not make an ambiguous cloud operation unambiguous by thinking harder.

The fine-tune that answered with last week’s policy

· 10 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerA fictional production incident about keeping an AI answer tied to current policy.

Ashvale Benefits (a fictional employee-benefits administration company) helps mid-sized employers manage health-plan enrollment, eligibility rules, and employee policy questions. It had eleven days before its busiest open-enrollment season.

The company had built an AI support assistant inside its member portal. The assistant was supposed to answer routine questions from each employer’s approved policy handbook, show the supporting section, and hand uncertain cases to the support team.

During open enrollment, that team expected eighty thousand questions. Most would be some version of the same thing: Does my plan cover this? When does coverage begin? Which form do I need?

The business goal was simple. Answer the routine questions immediately. Send the hard ones to a human. Never invent a benefit that did not exist.

There was one more rule from legal: every answer had to point to the policy section that supported it.

The retry that made the outage worse

· 11 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about retry ownership across an AI risk-decision path.

Kiteframe Pay (a fictional payment-risk infrastructure company) helps online marketplaces decide whether a card order should be approved, declined, or reviewed by a human. Merchants call Kiteframe during checkout, before they capture money or promise inventory to a shopper.

Most orders never need generative AI. Deterministic rules settle obvious cases in under 80 milliseconds. The difficult eight percent—new devices, unusual delivery patterns, sparse account histories—enter Aster, Kiteframe's AI-assisted risk analyst.

Aster was not a chatbot and could not approve a payment. It gathered approved evidence, used a language model to produce a typed risk recommendation with cited signals, and passed that recommendation to a deterministic merchant-policy engine. Without Aster, those ambiguous orders went to a manual-review queue. During a large sale, that queue could grow faster than Kiteframe's analysts could empty it.

At 02:13 on Tuesday, model latency rose sharply. Seven minutes later, Kiteframe's entire checkout success rate had fallen from 99.4% to 71%.

The model provider was recovering.

Kiteframe's retries would not let it.