Skip to main content

The right model that arrived too late

· 12 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyAdvancedA fictional production incident about correctly routed reasoning work that consumed its useful deadline.

IonWeave Semiconductor is a fictional chip manufacturer. Its fabrication plants run hundreds of tightly controlled tools that deposit, etch, measure, and clean material during chip production. A stopped tool is expensive. An unsafe restart is worse.

IonWeave built ToolSage, an internal AI assistant for line engineers. When a fabrication tool raised an alarm, ToolSage received a sealed machine-state snapshot, maintenance notes, and the approved runbook revision. It drafted a cited troubleshooting plan. It did not clear interlocks, change a recipe, or send commands to equipment. It had no equipment credentials.

Yuna Park (the inference reliability lead responsible for model routing, reasoning policy, and admission) owned the path from a ToolSage request to a completed model response.

A line engineer checked the response against the machine state shown by the tool. An equipment owner alone could authorize an approved runbook step or take manual control. The useful outcome was therefore not “model produced text.” It was “owner received enough verified information to decide safely while the incident was still actionable.”

How an alarm normally becomes a decision​

Loading the ToolSage journey…

The Tool Control Gateway sealed the alarm code, interlock bit, recent sensor state, maintenance notes, tool support status, recipe revision, and current runbook. ToolSage then applied route policy TS-31 to those authoritative fields.

Known routine alarms used Torch-Triage-12B/r2026.07.2, prompt template TRC-5, with no adaptive reasoning and at most 600 output tokens. Novel recipe alarms and active safety interlocks used Torch-Reason-70B/r2026.07.1, also with TRC-5, with an adaptive 800–8,000-token and at most 1,000 output tokens. Both ran on the fictional IonServe/r4.8.1 runtime.

TS-31 did not ask a model whether an interlock was active. It read interlock_active, supported-tool state, and recipe revision from the gateway and registry. Model utility came after hard eligibility. If a request was eligible, the selected model proposed a plan; the engineer verified it; the equipment owner decided what happened next.

That was the system Yuna took into a recipe ramp.

Concepts in this story5 concepts
Reasoning budgetGenerative AI · Inference policy

A limit or operating allocation for extra inference work assigned to a request. Its meaning and observability are release-specific, and a larger budget can improve some outcomes while increasing latency, cost, and interference with other requests.

Admission controlGenerative AI · Inference serving

The policy that decides whether estimated work may enter a finite serving system while preserving declared quality and deadline objectives. It needs a workload estimate; routing to a model does not itself reserve capacity.

Prefill versus decodeGenerative AI · Inference serving

Prefill processes the supplied input and may reuse matching prefix state; decode produces subsequent tokens over time. Healthy prefill or time-to-first-token metrics do not prove that long, variable decode work will finish within the useful deadline.

Deadline classGenerative AI · Inference serving

A workload class defined by the time remaining for its complete useful outcome, including required downstream work such as human verification. A priority label alone neither reserves the needed resources nor guarantees completion.

Qualified goodputGenerative AI · Inference economics

The rate of outputs that satisfy the declared quality, segment, and end-to-end timing contract. Raw tokens per second, utilization, first-token latency, or HTTP completion may all improve while qualified goodput falls.

The upgrade that raised plan acceptance​

During a ramp, new recipes naturally produced unfamiliar but low-consequence calibration alarms. The older fixed-budget configuration often returned a shallow checklist. The pinned Reason release used additional reasoning work when the snapshot was ambiguous. Line engineers accepted more of its plans without editing them.

The operational comparison looked strong. Line-engineer acceptance without editing rose. A repeated 9,000-token runbook prefix was increasingly reused. Worker prefill got faster. Expensive accelerators were busier. Projected spend still fit under the approved cap.

The route was not wrong. Interlocks were correctly sent to the Reason release. The demonstrated gain was higher line-engineer acceptance without editing in the dominant novel-routine segment. The failure lived between routing and useful completion.

The promise hidden inside ninety seconds​

For an active interlock, IonWeave's contract was end to end:

  1. admit the request or escalate early by t+10s;
  2. provide a complete cited plan by t+60s; and
  3. leave up to 30 seconds for owner verification and an action decision by t+90s.

That made the interlock request a : a request category whose useful service objective included its complete path and explicit failure behavior, not merely a queue priority flag.

The Reason pool had four decode lanes. Each lane processed one generation at a time. Work was non-preemptive: once a generation began, a later request could not reclaim that lane. The queue was FIFO.

ToolSage's reserved 2,000 reasoning tokens plus 600 output tokens for every Reason request. The estimate made sense for the pre-ramp distribution. But it was not a limit. The pinned release could use up to 8,000 reasoning tokens, and the scheduler allowed it to do so without revising the reservation.

The capacity ledger priced permission as an estimate.

Four lanes, eight long requests, one rare interlock​

At t=0, eight novel calibration requests—N1 through N8—arrived together. Each carried 10,000 input tokens, including a shared 9,000-token runbook prefix already present in the cache.

After a request acquired a lane, prefill took 0.6 seconds. Each novel request then used 7,200 reasoning tokens plus 600 visible output tokens. At the fictional measured rate of 200 generated tokens per second, the 7,800 generated tokens occupied a lane for another 39 seconds.

N1 through N4 took the four lanes. N5 through N8 waited.

At t=12, an active-interlock request, I1, arrived. It needed 4,000 reasoning tokens plus 600 output tokens: 23 seconds of decode after the same 0.6-second cached prefill.

Admission used the 2,000-token reasoning estimate for the running and queued work. It predicted I1 would complete at t=40.8, comfortably before I1's plan deadline at t=72.

Execution used the work the release was actually permitted to perform.

The first wave finished at t=39.6. FIFO then assigned all four lanes to N5 through N8. Those finished at t=79.2. Only then did I1 acquire a lane. Its cached prefill completed at t=79.8; decode completed at t=102.8.

I1 finished 90.8 seconds after arriving. It missed the complete-plan deadline by 30.8 seconds and the owner-decision deadline by 0.8 seconds. A plan arriving at the end of the owner window was not a successful response; there was no verification time left.

Loading the incident sequence…

The incident separates . Prefill processed the input and benefited from the cached prefix. Decode produced reasoning and answer tokens sequentially and still occupied scarce lanes for a variable duration. A healthy cache-hit rate and better time to first token did not imply timely completion.

The cache had done exactly what it promised. It avoided repeated prefix computation. It did not accelerate the 7,800 generated tokens or remove the FIFO wait in front of I1.

Why the dashboard stayed green​

The aggregate plan-acceptance measure was dominated by routine and novel-routine traffic. Line engineers accepted more of those plans without editing. Accelerator busy time rewarded a pool that was doing more work. Spend remained below its cap. Worker prefill p95 measured work after lane acquisition, not time waiting for a lane.

Even “high priority” was easy to overread. In this deployment, priority influenced a request attribute. It did not preempt running decode work, change FIFO ordering, reserve capacity, or prove deadline feasibility. The word described intent, not a schedulability guarantee.

The segment that mattered operationally was : the rate of responses that satisfied the required plan-acceptance, evidence, policy, and end-to-end timing conditions.

Only now did the complete comparison become diagnostic:

MeasureBeforeDuring ramp
Overall plan acceptance88.9%92.7%
Reason-pool prefix-cache hit rate82%91%
Worker-side prefill p951.3s0.8s
Accelerator busy time72%89%
Projected spend versus cap88%96%
Safety-interlock qualified goodput92%54%

The first five rows reinforced the rollout case. The last row represented just 0.4% of the ramp window—SAFETY_INTERLOCK_DURING_RECIPE_RAMP—and fell even while aggregate acceptance and hardware utilization rose.

Yuna reconstructed the path request by request:

authoritative class
→ TS-31 route
→ admitted reasoning ceiling
→ lane acquisition
→ cached prefill
→ variable decode
→ complete cited plan
→ engineer verification
→ owner decision

The first broken contract appeared at admission. The system reserved less work than it allowed the model to perform. FIFO scheduling then let a common novelty burst consume every lane ahead of a rare deadline class.

The tempting capacity answer​

The release council proposed enabling adaptive reasoning across the fleet and adding replicas. The first part extended the documented line-engineer acceptance gain. The second could reduce latency for a measured arrival rate.

But neither changed the unenforced 2,000-token estimate or the FIFO contract.

Twenty-five percent more capacity could improve the observed ramp. It could not establish the interlock deadline bound while each accepted request was allowed to consume far more work than admission reserved. A sufficiently broad novelty burst could fill the additional lanes with the same underpriced, non-preemptive jobs.

That did not mean scaling was useless. Once work ceilings and deadline feasibility were explicit, capacity planning could answer a well-formed question: how much admitted demand should this design sustain? Before that correction, replicas made the unsafe region harder to locate rather than impossible to enter.

Two corrections, two different costs​

Yuna's team kept deterministic interlock authority and the same pinned model releases. They compared two production designs.

Design A: protect one deadline lane​

One of the four Reason lanes became interlock-only. Novel routine requests received an enforceable ceiling of 4,000 reasoning plus 600 output tokens. Interlocks received 6,000 reasoning plus 600 output. Admission reserved the full class ceiling and escalated by t+10 when the protected lane could not meet t+60.

In the worked cohort, I1 began immediately after arrival and finished 23.6 seconds later.

The price was visible. The protected lane could idle while novel work waited. General Reason capacity fell from four lanes to three. The novel-routine cap could reduce line-engineer acceptance without editing on hard cases. An interlock burst could still exceed the protected lane, so some requests would abstain rather than receive an automatic plan.

Design B: pool the lanes and admit the whole path​

All four lanes remained pooled. Before inference, deterministic checks rejected missing alarm fields, stale sensor state, unsupported recipe revisions, or absent runbooks. The route still came from TS-31; Torch-Triage did not judge interlock risk.

Admission reserved the enforceable class ceiling and used a conservative measured service rate to test whether accepting a request would preserve existing plan deadlines. Ready work was ordered by earliest plan deadline first. Execution remained non-preemptive, so the feasibility check included possible blocking by work already running. Infeasible requests escalated early.

In the worked cohort, I1 finished 35.2 seconds after arrival. Under its full 6,600-token ceiling, the bound was 45.2 seconds—still inside the plan deadline.

Pooling used capacity more flexibly than a permanently protected lane, but it was not free. Conservative reservations reduced automated volume. Enforced caps could lower line-engineer acceptance without editing. Running jobs could still block an urgent arrival. Larger bursts produced earlier abstentions, moving cost to human operations.

Neither design was universally better. Design A bought the cleanest rare-class isolation and paid with idle reserve. Design B bought pooling efficiency and paid with conservative admission, scheduler complexity, and more explicit abstention. Both made the full work contract enforceable before adding replicas.

Loading the corrected topology…

The topology also preserves the most important non-model boundary. The gateway and registry own equipment facts. ToolSage owns policy and proposal orchestration. The inference control plane owns budget and scheduling. Humans own verification and action. There is no edge from a model pool to fabrication equipment.

The review question that survived the incident​

A month later, another team proposed sharing a reasoning-heavy code-analysis pool with emergency rollback requests. The normal code reviews had flexible deadlines and benefited from larger reasoning budgets. Rollback requests were rare, but a correct patch after the rollback window was operationally useless.

Yuna did not begin by asking whether the router selected the most capable model. She asked for five artifacts:

  1. the authoritative request classes and hard eligibility rules;
  2. enforceable reasoning and output ceilings by class;
  3. the separate prefill, queue, and decode service distributions;
  4. the complete-path deadline, including human verification; and
  5. qualified goodput by the rare deadline class.

Then she asked what 25% more capacity could prove. The answer was deliberately narrow: it could improve performance at a measured workload. It could not replace a bound on admitted work or a policy for early failure.

The transferable rule was not “always reserve a lane” or “always use EDF.” It was:

Route correctness is only one link. Price the work the model may actually perform, schedule it against the deadline that makes the answer useful, and measure success after the final required human step.

Concepts in this story5 concepts
Reasoning budgetGenerative AI · Inference policy

A limit or operating allocation for extra inference work assigned to a request. Its meaning and observability are release-specific, and a larger budget can improve some outcomes while increasing latency, cost, and interference with other requests.

Admission controlGenerative AI · Inference serving

The policy that decides whether estimated work may enter a finite serving system while preserving declared quality and deadline objectives. It needs a workload estimate; routing to a model does not itself reserve capacity.

Prefill versus decodeGenerative AI · Inference serving

Prefill processes the supplied input and may reuse matching prefix state; decode produces subsequent tokens over time. Healthy prefill or time-to-first-token metrics do not prove that long, variable decode work will finish within the useful deadline.

Deadline classGenerative AI · Inference serving

A workload class defined by the time remaining for its complete useful outcome, including required downstream work such as human verification. A priority label alone neither reserves the needed resources nor guarantees completion.

Qualified goodputGenerative AI · Inference economics

The rate of outputs that satisfy the declared quality, segment, and end-to-end timing contract. Raw tokens per second, utilization, first-token latency, or HTTP completion may all improve while qualified goodput falls.


IonWeave Semiconductor, ToolSage, its people, releases, runtime, route policy, incident, workload, token rates, and metrics are fictional. Technical grounding is bounded to mechanism claims: Anthropic's extended-thinking documentation documents release-specific reasoning controls and their latency/cache interactions; vLLM automatic prefix caching explains that prefix reuse avoids repeated prefill computation but does not accelerate decode; and vLLM disaggregated prefilling distinguishes prefill from decode and cautions against assuming a universal throughput improvement. These sources do not establish IonWeave's fictional semiconductor workflow, safety policy, token rates, deadline, or incident metrics.