Timeout
A limit on how long a caller will wait before treating an operation as failed. A timeout ends the wait; it only stops the underlying work when cancellation is propagated.
Definitions are grouped into reader-facing learning areas and linked to the stories where each concept becomes load-bearing.
A limit on how long a caller will wait before treating an operation as failed. A timeout ends the wait; it only stops the underlying work when cancellation is propagated.
A new attempt after a failed one. Retries help with transient faults, but they add load and can repeat side effects unless the operation is safe to retry.
A retry schedule in which the delay grows exponentially after each failure, usually up to a cap. It reduces pressure on a dependency that is still unhealthy.
Random variation added to retry delays so many clients do not retry in lockstep. It spreads recovery traffic over time instead of creating another spike.
A feedback loop where failed requests trigger enough retries to raise load, deepen the failure, and delay recovery.
A stateful guard that stops calls to a failing dependency for a limited time, then probes for recovery. It protects both systems; it is not another retry mechanism.
Closed allows calls and records outcomes. Open rejects calls immediately. After a cooldown, half-open admits a small number of probes and uses their results to close or reopen the circuit.
A deliberately reduced result returned when the preferred dependency or path is unavailable—for example, cached data without personalization.
Parametric knowledge is encoded implicitly in a model’s learned weights. Non-parametric knowledge lives outside the model—such as in documents or an index—and is fetched when needed. External knowledge can be revised without retraining the model.
A generation pattern that retrieves relevant external information for the current query and includes it in the model’s input before the answer is generated. Retrieval changes the context, not the model’s weights.
RAG supplies external knowledge at inference time; fine-tuning changes model weights through training. Use retrieval as the default for changing, sourceable facts and fine-tuning for learned behavior, format, or task adaptation. They can be combined.
The indexing pipeline prepares source content before questions arrive: load, split, enrich, and store it. The query pipeline runs per request: interpret the question, retrieve evidence, assemble context, generate, and validate the answer.
Connecting generated claims to supplied evidence or an authoritative external source. Grounding makes answers inspectable and can reduce unsupported claims, but retrieval and generation can still fail.
Persisting workflow progress so a long-running agent can resume after waiting or process loss. A checkpoint records completed orchestration state; it does not prove that an unacknowledged external effect failed or that the outside world is still unchanged.
A stable caller-provided identifier for one logical operation. Repeating the same request with the same key lets the receiver suppress an additional effect within its documented scope and retention window; a new key represents new intent.
A guarantee that one logical operation is applied once inside a stated boundary, usually by combining atomic state changes, deduplication, or idempotency. A workflow or broker guarantee does not automatically make a separate external side effect exactly once.
A monotonically increasing epoch attached to work on a protected resource. The resource or gateway rejects requests carrying an older epoch, preventing a stale lease holder from beginning new effects after a successor takes over.
End-state evaluation checks where the task finished. Trajectory evaluation also checks the actions, observations, policy decisions, costs, and prohibited intermediate effects used to get there. A correct final state can follow an unacceptable path.
The broader capability created when several individually permitted operations can be chained. Local least-privilege checks do not establish that the complete path or resulting effect is authorized.
An authorization decision over the complete path from protected input through transformations to a consequential destination. It checks whether this data may reach this sink for the current purpose, not only whether each intermediate tool call is allowed.
The supported relationship between a versioned record and the person or entity it concerns. A model may propose this relationship from semantic evidence, but release authority requires corroborating source metadata, policy, or explicit review.
Human authorization tied to one canonical proposed effect: its target, parameters, evidence, constraints, destination, and expiry. A material change creates a new approval question instead of inheriting consent from the surrounding task.
Re-evaluating current principal, workload, resource, purpose, policy, approval, and effect immediately before a consequential action. Permission observed during planning or approval is historical evidence, not automatic present authority.
A versioned rule that maps model evidence, thresholds, current context, capacity, and constraints into an action such as automate, inspect, review, defer, or abstain. The model supplies evidence; the policy owns the action boundary.
Running a candidate model beside the live system without letting it control production actions. This limits exposure, but observed outcomes still reflect the incumbent policy unless the evaluation creates an independent evidence path.
Labels observed for a non-random subset because a prior decision policy determines which cases receive the action, inspection, approval, review, or follow-up that can reveal an outcome.
The explicit register of eligible units available for selection into a sample. It is distinct from the target population, the design that chooses units from the frame, and the observed cohort that actually produces usable labels.
A bounded live policy that deliberately acquires evidence where an incumbent and candidate would choose different actions. It can estimate value in that disagreement region, but it does not establish whole-population model quality.
An outcome is the real condition the system wants to know, an observation is evidence produced by a named method and authority, and a label is the versioned value recorded for learning. An action or proxy may correlate with the outcome without establishing a label.
An unknown label means the declared authority did not establish the outcome. A negative label means authoritative observation supports the negative class; a missing record, elapsed deadline, release, or lack of complaint cannot silently convert unknown into negative.
The time authoritative label evidence became usable by an evaluation or training snapshot. It can differ from when the real outcome existed and when the observation occurred, so a historical dataset must use only the label revision available at its declared cutoff.
The trace from a label to its unit, source, observation method, authority, timestamps, adjudication, and revision, together with the datasets and model decisions that consumed each version. A correction creates attributable new evidence rather than invisibly rewriting history.
Constraining a model response to an allowed shape, grammar, or schema while tokens are generated. It can guarantee representation and permitted members; it does not establish that the resulting values are true, compatible, or authorized for use.
Conformance to a declared representation contract such as required fields, types, cardinality, and enum membership. Structural validity says the artifact is well formed under that contract, not that its values form a valid business decision.
A relationship established by authoritative rules or records showing that individually valid values can be used together for the current request. It must be checked at the boundary that can still prevent the effect.
An explicit outcome that declines to produce or accept an automated answer when the system cannot satisfy its evidence, capability, risk, or deadline contract. Abstention must lead to a defined hold, fallback, or human path.
A limit or operating allocation for extra inference work assigned to a request. Its meaning and observability are release-specific, and a larger budget can improve some outcomes while increasing latency, cost, and interference with other requests.
The policy that decides whether estimated work may enter a finite serving system while preserving declared quality and deadline objectives. It needs a workload estimate; routing to a model does not itself reserve capacity.
Prefill processes the supplied input and may reuse matching prefix state; decode produces subsequent tokens over time. Healthy prefill or time-to-first-token metrics do not prove that long, variable decode work will finish within the useful deadline.
A workload class defined by the time remaining for its complete useful outcome, including required downstream work such as human verification. A priority label alone neither reserves the needed resources nor guarantees completion.
The rate of outputs that satisfy the declared quality, segment, and end-to-end timing contract. Raw tokens per second, utilization, first-token latency, or HTTP completion may all improve while qualified goodput falls.
The complete compatibility surface that jointly determines a generative attempt’s behavior: model, tokenizer, template, adapter, context and reasoning policy, decoding, validator, runtime, precision, topology, provider, route policy, and evaluation release where applicable. Model weights alone do not identify it.
An application-owned immutable identifier that resolves to one behavior-bundle manifest before an attempt runs. Friendly aliases and requested or returned provider model names are useful evidence, but none is a complete or stable release identity by itself.
The supported period in which more than one behavior release or release-produced representation remains active. Equal schemas or locally valid releases do not prove that outputs produced under different semantic contracts can safely be combined.
The parentage connecting a request and each retry, fallback, or continuation to its exact behavior release, input artifacts, accepted output, authority receipts, and durable descendants. Human edits may add ancestry; they do not erase the earlier influence.
Evidence that every declared recovery plane has reached its intended boundary: future routing, in-flight work, cached or generated representations, committed application state, external artifacts, evaluation rows, and learning data. Restoring future traffic is one plane, not proof that rollback is complete.
A test whose expected value was recorded from the system’s own observed output so that later change becomes visible. It documents what the code currently does, which is not the same as what the code should do: a pass establishes agreement with an earlier observation, not conformance to a rule.
Where a test’s expected value came from — a stated rule, policy, or invariant, or an observation of the system itself. Provenance decides what a passing test can prove, and an expectation recorded from output cannot separate a change made for the intended reason from one made for an unintended reason.
A statement of intended behavior written so that a machine can check it and the rule’s owner can review it — an invariant over all inputs, or example rows authored from the rule text rather than generated from the running system. Its value depends on being authored independently of the code it checks.
A business rule published with a version and a validity window, so behavior can change on a date without erasing what was correct before. Effective dating makes a behavior change reviewable and reversible, and it obliges the system to keep older versions computable for as long as re-pricing, disputes, or audits can reach back.
Resolving a value that lives outside the event using a stated time rather than whatever the value happens to be now. The stated time can mean two different things — the time the event describes, or a declared cut in when the system asserted the value — and a design that supplies only the first still returns today’s assertion about a past day.
The property that re-running a computation over the same input produces the same output, where the input includes every external value the computation reads. Reproducing offsets, keys, order, event timestamps, watermark progression, and code version is necessary and not sufficient; anything resolved at processing time is different on every run by construction.
Valid time is the period during which a fact is true of the world. Assertion time is the period during which a system held that description of it. A source recording only the current description can be correct at two different moments and give two different answers, and a rebuild that does not say which assertion it wants silently takes the newest one.
A derived value or constraint whose scope spans more than one ordering key. Per-key ordering and per-key replay scope are both defined below its granularity, so neither preserves it: a rebuild covering only some of the invariant’s keys can produce a result that no consistent rebuild of any single version would produce.
The point after which a derived record stops being a computation and becomes a commitment. Before it, a change is a correction; after it, the same change is a restatement that needs a scope, an owner, and an authorisation. Finality is a domain rule rather than a system property, and different boundaries have different owners and different reopening costs — some belong to another party and cannot be reopened at all.
The machinery that changes a system: creating, describing, updating, deleting, and listing resources, and propagating those changes to wherever they take effect. Placement, provisioning, migration, and credential issuance are control-plane work.
The daily business of serving requests from state a component already holds. A data plane is usually simpler than the control plane that arranges it, which is why it can stay healthy while the control plane is degraded.
A complete, independent instance of a workload that shares no state with other cells and serves a subset of customers or resources. Cells narrow the scope of failures that happen inside a cell; they do not isolate a dependency every cell’s traffic must cross.
The property that a system keeps working correctly while a dependency is impaired, because nothing has to change for it to keep working. It is bought with pre-provisioned capacity or pre-published state, and paid for in staleness, cost, or operational duplication.
One region of state that a single cancellation actually reaches, together with the system that owns it. A live interaction has several — audio playout, model generation, turn and command intent, the tool request, and the external effect — and each acknowledges separately, at a different speed, with a different meaning. How fast an acknowledgement arrives tells you nothing about what it proves, and the domain that owns the effect may never acknowledge at all.
The point at which a system decides a person has finished speaking and treats their words as an input it will act on. Silence thresholds and semantic endpointers make that decision probabilistic, so a committed turn is a guess about a human thought, not a fact about it — and an irreversible effect must not depend on one.
The boundary between output a system generated or sent and output a person actually heard. Realtime models emit audio faster than realtime, so an interruption leaves buffered speech that was delivered and never played. Conversation history must be reconciled to the played position, or a resumed session will confidently refer to a confirmation nobody received.
The moment in an effect’s lifecycle after which cancellation no longer exists, and the only remaining option is a separate compensating request that another party may refuse. Before it, cancellation can be exact because no external system has been told anything; after it, a cancellation is a negotiation. Holding an effect moves this point later; it never removes it.
Dropping a late asynchronous result from the current conversation because its turn, topic, or command has moved on. It is the right conversational policy and the wrong records policy: a receipt must always enter the authoritative effect log, while whether it is spoken aloud is a separate decision with a separate owner.