Skip to main content

The package no technician had inspected

· 14 min read
Fault Lines Editorial
Fictional incidents. Exact technical vocabulary.
Architecture storyBeginnerPackProof inspected six camera views. The missing evidence was an independent physical inspection of the package.

MorrowVale Foods is a fictional refrigerated-food manufacturer. Its packaging lines print and apply the required label to each sealed retail package before cases move to the warehouse. A wrong artwork revision, unreadable allergen text, missing language panel, or damaged lot code can place that package on hold.

MorrowVale built PackProof, a fictional multimodal inspection service, to help its quality team find those problems. Six synchronized cameras photographed every visible side of one sealed package. PackProof read those images beside the package's approved artwork specification and proposed an anomaly score, evidence quality, suspected issue types, and a priority for physical inspection.

PackProof inspected every package's six camera views. What did not happen for most low-risk packages was the independent physical inspection that MorrowVale had named as its authoritative evidence.

Leena Rao, the quality systems engineer responsible for PackProof's inspection evidence and training pipeline, had kept the model's authority deliberately narrow. PackProof could propose where scarce technician attention should go. A deterministic inspection-policy service chose the route. Warehouse systems owned release and hold. Only a trained quality technician using MorrowVale's fictional QL-4 protocol could create a physical pass or fail observation.

That was the intended design. The failure appeared later, where evidence became training data.

How one package normally moved​

Loading the package journey…

A line operator first attached a unique package_id to the sealed retail package and the immutable artwork_spec_id active for its SKU and production lot. The six cameras then captured the same unit at captured_at, along with line, shift, printer, and lighting telemetry.

PackProof's result was a proposal, not a compliance decision. The deterministic policy applied four ordered rules:

invalid capture -> HOLD_CAPTURE
degraded evidence -> DIVERT_PHYSICAL
anomaly score at or above threshold -> DIVERT_PHYSICAL
predeclared rotating-observation sample -> DIVERT_PHYSICAL
otherwise -> RELEASE_ORDINARY

When selected, a quality technician scanned that same package and checked it under standardized light against the bound artwork specification. The fictional QL-4 protocol had to finish within 30 minutes of camera capture and before rework, relabelling, or shipment could change the object under judgment. The technician did not see PackProof's suspected issue list.

The Quality Inspection Registry then recorded one of five reader-visible states:

PENDING_SAMPLE
PHYSICAL_PASS
PHYSICAL_FAIL
INCONCLUSIVE
NOT_OBSERVED(
reason = NOT_SELECTED | WINDOW_EXPIRED | PACKAGE_UNAVAILABLE
)

PHYSICAL_FAIL could trigger a bounded pallet hold. RELEASE_ORDINARY was also a legitimate operational outcome. Neither warehouse action was allowed to rewrite the physical-evidence state.

The weekly query that made the dataset look complete​

Operations legitimately needed terminal warehouse records. Seven days after shipment, most packages had no open task and no complaint. The model-training team also wanted a Boolean target with enough rows for weekly retraining.

The projection that joined those needs looked competent:

case
when q.observation_state = 'PHYSICAL_FAIL' then 1
when q.observation_state = 'PHYSICAL_PASS' then 0
when w.shipped_at < :cutoff
and c.first_complaint_at is null then 0 -- unsafe
else null
end as label_has_required_label_defect

The first two branches preserved QL-4 evidence. The third manufactured evidence. It called a package compliant because it had shipped and no complaint had reached MorrowVale by the cutoff.

That proxy was attractive for three reasons. It closed the weekly cohort. It produced far more negatives than QL-4 alone. And the resulting dashboard had no awkward unknown column.

But a customer complaint depended on shipment, noticing a problem, deciding to report it, identifying the package, and the report reaching the right system. Complaint absence was a real record of no complaint received. It was not a technician's physical reading of the label.

The query did not merely lose nuance. It changed what 0 meant from “QL-4 physically passed this package” to “QL-4 passed it, or the organization did not hear otherwise.”

The slice PackProof learned not to inspect​

On a Tuesday maintenance window, MorrowVale replaced one lighting controller. The new controller created a narrow glare pattern on one artwork family. The cameras still returned valid images, but PackProof confidently low-scored some real printing defects in that lighting-and-artwork slice.

Those low scores had two effects.

First, fewer affected packages entered the ordinary QL-4 path. The registry honestly recorded most as NOT_OBSERVED.

Second, the weekly query discarded that honesty. When those packages shipped without a complaint, it stored 0. Ordinary retraining then consumed the invented negatives. The next PackProof release became still more confident in that slice and diverted fewer packages.

The loop was not “the model trained on its own predictions.” It was narrower and more dangerous:

model output changes who receives physical inspection
-> missing observations are replaced by an operational proxy
-> the proxy is stored as a negative training label
-> retraining reinforces the under-observed slice
-> the model sends even fewer packages for observation there

No shadow candidate or model tournament was involved. This was the normal weekly release learning from a label table whose meaning had drifted.

The audit that could make only a bounded claim​

MorrowVale also ran a small rotating quality audit. Its membership was locked from package_id, line, shift, artwork, and lighting stratum before PackProof's output or any technician result was known. Those non-model strata determined membership. After PackProof ran, the audit report attached the package's current score band for analysis; that post-hoc field could not change whether the package had been selected. A package selected by both the model route and the audit received one QL-4 inspection and retained both route reasons.

During the next scheduled audit, technicians recorded multiple PHYSICAL_FAIL results among sampled packages from the new glare slice. Warehouse operations placed the affected pallet on hold and began bounded reinspection.

That result did not prove a plant-wide defect rate, a fleet-wide model error rate, or a public-safety outcome. It proved that those physically inspected packages failed QL-4. It also gave Leena a concrete trace to follow.

The trace showed comparable packages with this sequence:

Quality Inspection Registry: NOT_OBSERVED(NOT_SELECTED)
warehouse: shipped
complaints: none received by weekly cutoff
training table: 0, interpreted as compliant

Reality had not labelled those packages negative. The dataset contract had.

Loading the incident sequence…

Four different records about one package​

The audit trace made it possible to name the four things that MorrowVale's weekly table had collapsed into one Boolean column.

QuestionMorrowVale recordExample
What was true of the physical package?Real-world outcomeIts printed label conformed to the bound artwork specification, or it did not.
What did the named protocol actually establish?QL-4 observationPHYSICAL_PASS, PHYSICAL_FAIL, or INCONCLUSIVE.
What value may a learning job consume?Training label0, 1, or NULL, derived under an explicit rule.
What else happened operationally?Proxy or action recordReleased, shipped, no complaint, no return, or placed on hold.

That separation is . The physical package had an outcome whether or not MorrowVale learned it. QL-4 could observe that outcome within a bounded protocol. A label was the value a later dataset stored after applying a projection to the observation. Shipment and complaint absence described different events.

For high-scoring packages, the distinction often looked academic because a technician usually produced a timely QL-4 result. Low-scoring packages exposed it. If policy did not request QL-4, the registry immediately recorded NOT_OBSERVED(reason=NOT_SELECTED). That was not a failed inspection and not a pass. It meant the named authority had produced no physical evidence for that package.

This is . A negative label says the authoritative observation established the target was absent: here, QL-4 found no required-label defect. Unknown says that evidence does not exist or did not resolve the question. Unknown is not a softer form of negative.

Two clocks, not one timestamp​

Leena first stopped the affected training snapshots from feeding another release. Then she examined a less visible problem: even genuine QL-4 evidence did not become usable at the instant it was observed.

The registry kept two clocks:

  • observed_at: when the technician completed the physical protocol on the package;
  • label_available_at: when the signed observation revision became available to downstream dataset builders.

The second is . It prevents a historical training snapshot from using evidence that existed in the physical world but had not yet reached the label system at that snapshot's cutoff.

Suppose a technician completed QL-4 at 10:18, but a network interruption delayed the signed registry entry until 10:43. A dataset with a 10:30 availability cutoff must not claim it consumed that result. A later reconciled snapshot may include it, with a new identity.

The distinction also made late work explicit. A QL-4 result after the 30-minute target window did not silently overwrite NOT_OBSERVED(reason=WINDOW_EXPIRED). It became a new observation record with its own stated scope.

The corrected label boundary​

The unsafe branch disappeared. The training path now asked only the Quality Inspection Registry for the final eligible QL-4 observation under the bound package and artwork specification.

type QL4Observation = {
observationId: string;
packageId: string;
artworkSpecId: string;
capturedAt: string;
observedAt: string | null;
labelAvailableAt: string;
labelRevision: number;
state:
| 'PENDING_SAMPLE'
| 'PHYSICAL_PASS'
| 'PHYSICAL_FAIL'
| 'INCONCLUSIVE'
| 'NOT_OBSERVED';
};

function toBinaryTrainingLabel(observation: QL4Observation): 0 | 1 | null {
if (observation.state === 'PHYSICAL_FAIL') return 1;
if (observation.state === 'PHYSICAL_PASS') return 0;
return null;
}

const observation = qualityInspectionRegistry.getFinalQl4Observation({
packageId: capture.packageId,
artworkSpecId: capture.artworkSpecId,
capturedAt: capture.capturedAt,
targetWindowMinutes: 30,
availableAtOrBefore: snapshot.labelCutoff,
});

const row = {
packageId: capture.packageId,
artworkSpecId: capture.artworkSpecId,
observationId: observation.observationId,
observedAt: observation.observedAt,
labelAvailableAt: observation.labelAvailableAt,
labelRevision: observation.labelRevision,
target: toBinaryTrainingLabel(observation),
};

The important property is not TypeScript's null. It is the authority check around the function: same package, same artwork specification, explicit observation window, availability cutoff, and exact registry revision.

That record establishes : the ability to trace one training value to the observation, identity, timing, rule, and append-only revision that produced it. If an authorized adjudication later revised INCONCLUSIVE to PHYSICAL_FAIL, the registry appended a new revision. It did not mutate an old snapshot invisibly. The catalog identified and invalidated every affected snapshot.

The corrected state machine prevented fabricated labels. It did not acquire missing physical evidence.

The 10,000-package cohort after the correction​

Leena ran a fictional 10,000-package fixture through the corrected projection. These values exist only to make the state arithmetic inspectable.

State at observation-window closePackagesPositive 1Negative 0Unknown / excluded
PHYSICAL_FAIL15015000
PHYSICAL_PASS79007900
INCONCLUSIVE200020
Requested but unavailable or expired400040
Not selected for QL-49,000009,000
Total10,0001507909,060

The binary cohort contained 940 packages, not 10,000. The other 9,060 were not counted as failures, but they were not declared compliant either.

That smaller number made some model dashboards look worse. It made the dataset's claim much more precise: every binary target traced to a completed, eligible physical observation.

Concepts in this story4 concepts
Outcome, observation, and labelML systems · Label contracts

An outcome is the real condition the system wants to know, an observation is evidence produced by a named method and authority, and a label is the versioned value recorded for learning. An action or proxy may correlate with the outcome without establishing a label.

Unknown versus negative labelML systems · Label contracts

An unknown label means the declared authority did not establish the outcome. A negative label means authoritative observation supports the negative class; a missing record, elapsed deadline, release, or lack of complaint cannot silently convert unknown into negative.

Label availability timeML systems · Label timing

The time authoritative label evidence became usable by an evaluation or training snapshot. It can differ from when the real outcome existed and when the observation occurred, so a historical dataset must use only the label revision available at its declared cutoff.

Label lineage and revisionML systems · Label provenance

The trace from a label to its unit, source, observation method, authority, timestamps, adjudication, and revision, together with the datasets and model decisions that consumed each version. A correction creates attributable new evidence rather than invisibly rewriting history.

Design A: authority-first labels with reconciliation​

MorrowVale's minimum correction kept the Quality Inspection Registry as the sole label authority.

Every requested inspection began as PENDING_SAMPLE. Every final state stayed explicit. Only final in-window PHYSICAL_PASS and PHYSICAL_FAIL results entered the binary cohort. Snapshot records bound both clocks and the exact label_revision. An append-only revision caused affected snapshots to be invalidated and rebuilt under a new identity.

This design bought truth eligibility, authority, and reproducibility. It did not buy representative coverage. PackProof still influenced which packages received ordinary QL-4 inspection.

The costs were not hidden: a smaller binary cohort, slower learning, delayed reconciliation, and many unresolved packages. The team had to operate revision tracking and snapshot invalidation instead of treating a weekly table as final.

Design B: add protected rotating observation​

Where that missing coverage mattered, MorrowVale extended Design A with the rotating observation program.

Membership was locked before model output and technician result from package_id, line, shift, artwork, and lighting stratum—never from the current model score. After PackProof produced its output, MorrowVale attached the current score band to the audit record for post-hoc coverage reporting. That report did not alter membership. The protected route was unioned with the ordinary risk route, so PackProof remained useful for focusing inspection while the protected path acquired some evidence outside its attention.

The design reported selected packages that could not be inspected. WINDOW_EXPIRED and PACKAGE_UNAVAILABLE remained unknown. The team did not quietly replace them or extrapolate from completed inspections as if every selected package had complied.

This extension bought evidence in under-observed slices and a way to notice when the model's ordinary route stopped looking somewhere. It did not buy certainty, a universally unbiased dataset, or a population defect rate.

It also spent real capacity: technician time, packages held long enough for QL-4, slower line release, fewer inspection slots for packages already predicted high-risk, and sampling uncertainty. Quality and operations had to decide how much protected observation they could afford.

Loading the production topology…

The two designs were not rivals. Design B depended on Design A's label authority and lineage. Without that foundation, more inspections would still flow into an ambiguous label table.

Transfer the contract to a pump​

A predictive-maintenance model sends high-risk pumps for teardown. Low-risk pumps keep operating and therefore have no teardown record. Thirty days later, a training query labels every pump without a fault code as NO_FAULT.

Before accepting that table, identify:

  1. the prediction unit, operational action, physical outcome, observation authority, and stored training label;
  2. which pump records are true negatives, pending, unknown, or inconclusive;
  3. why adding an explicit unknown state repairs the dataset's claim but does not create teardown evidence; and
  4. how a protected observation path could acquire evidence outside the model-selected route, and what teardown capacity, downtime, and equipment costs it would consume.

If “no fault code” can be produced simply because no teardown occurred, it is an operational record—not yet a negative outcome label.

Leena's final rule was equally short:

A decision can explain why evidence is missing. It cannot turn the missing evidence into a result.

Evidence and fiction note​

MorrowVale Foods, PackProof, QL-4, Leena, the lighting change, identifiers, package counts, audit, pallet hold, and all operational details are fictional.

The technical mechanism is grounded in current and durable research. A 2026 ICLR paper on positive-unlabeled evaluation analyzes evaluation where negatives are not directly observed. A 2026 study of deployed label feedback documents how deployment feedback can affect labels. The June 2026 Information Systems paper HILTS: Human–LLM collaboration for effective data labeling proposes and evaluates a framework and interactive system that combine LLM pseudo-labeling with targeted human review; it supports the narrower point that sampling, machine-produced labels, human correction, and validation design form one evidence system. A 2026 industrial label-inspection paper supports the plausibility of automated visual label inspection. The older selective-labels foundation is cited as durable background, not as a current landscape signal.

None of those sources documents this fictional company or incident, and this story does not claim that one bounded audit estimates a plant-wide defect or safety rate.