Skip to main content

Engineering guide / model releases

Release a tested system, then watch the work it changes.

A reachable prediction endpoint proves that something is running. Production readiness also depends on the inputs it receives, the decisions it influences, the evidence available when behavior changes, and the route back to a known operating state. Build those controls into the release.

The release boundary

Keep model evidence attached to the configuration users actually receive.

A production model release combines a model or provider configuration with preprocessing, application code, decision thresholds, dependencies, access rules, and an operating contract. Validate the assembled path in its target environment. Changing one of these parts can change behavior even when the model name stays the same.

Offline evaluation
Tests an identified candidate against an identified dataset and scoring method. Check leakage, relevant case coverage, and the ordinary workflow baseline. Record what the evaluation could not establish.
Serving acceptance
Tests input shape, feature computation, output handling, authorization, workload, and dependency failure. A model that performed well in a notebook can receive different features in a service.
Live acceptance
Observes the released workflow under permitted exposure. Request health, model quality, and operational effects need distinct measures and owners; none should silently stand in for another.

Four release contracts

Make the deployment review concrete.

Use an internal queue classifier as an example. It suggests a category for incoming work while people retain routing authority. A new release may change the model, preprocessing, or threshold. The contracts below explain what must remain inspectable as that release moves toward use.

Identity and evaluation

01

Release question: Is this the system that passed the review?

Required artifact
A release manifest linking model version, code, preprocessing, configuration, dependencies, and evaluation evidence.
Inputs
Candidate artifacts, dataset version, scoring rules, baseline results, case coverage, and known limits.
Implementation
Build a versioned release and record its artifact identity. Test the assembled request path rather than treating the model file as the complete product.
Accountable owner
The release owner accepts the permitted scope and unresolved limits.
Acceptance evidence
The candidate in the target environment produces expected results on agreed regression cases.
Failure to test
A threshold or preprocessing change reaches users under the old evaluation label.
Operational record
An inspectable manifest and a record of who approved which release for which use.

Serving and capacity

02

Release question: How does a request become a usable result?

Required artifact
An input/output contract, workload envelope, timeout policy, and result-delivery path.
Inputs
Request sizes, concurrency, freshness needs, feature availability, dependency quotas, and downstream consumers.
Implementation
Choose batch for scheduled work, synchronous serving for bounded interactive waits, or queued jobs when work must outlive a request. Validate inputs and distinguish rejection, failure, and pending work.
Accountable owner
Platform and product owners agree the supported workload and user-visible fallback.
Acceptance evidence
Representative load and fault tests include cold starts, unavailable features, saturation, and recovery.
Failure to test
A client retry starts another expensive job or a stale feature is silently accepted.
Operational record
Measured service behavior, resource limits, and a defined response when capacity is unavailable.

Monitoring and interpretation

03

Release question: Which signal should cause which action?

Required artifact
A monitoring plan separating service health, input changes, model outcomes, and workflow effects.
Inputs
Release identity, request cohorts, latency and error data, approved samples, and outcome labels when available.
Implementation
Break down observations by release and relevant task group. Track missing features and input changes separately from measured prediction quality. Join later outcomes to the correct prediction identity.
Accountable owner
Model and workflow owners interpret quality evidence; operators respond to service incidents.
Acceptance evidence
A deliberately introduced bad case triggers the intended alert and reaches an owner who can act.
Failure to test
An overall average conceals a failing category, or delayed labels create a false impression that quality is already known.
Operational record
Alert definitions, evidence windows, label coverage, escalation, and unresolved monitoring limits.

Recovery and ownership

04

Release question: What happens to users and unfinished work if this release is withdrawn?

Required artifact
A tested plan for stopping exposure, restoring or bypassing inference, and reconciling affected work.
Inputs
Previous release compatibility, queue state, stored outputs, caches, downstream effects, and incident responsibilities.
Implementation
Exercise recovery with representative pending and completed requests. Version result records so outputs from a withdrawn release remain identifiable.
Accountable owner
A named operator can stop exposure; the workflow owner decides how affected work is corrected.
Acceptance evidence
The ordinary process remains usable after withdrawal and unresolved items have an explicit destination.
Failure to test
Traffic returns to the old model while incorrect stored classifications continue influencing work.
Operational record
A recovery rehearsal, affected-record query, owner roster, and retirement procedure.

Worked release decision

Healthy requests can still produce the wrong workflow.

Imagine release R-12 of the queue classifier. Its API remains responsive, but one category sends substantially more items to manual review than the accepted baseline. These are hypothetical observations. The point is to predefine how evidence changes exposure, not to copy a universal threshold or rollout percentage.

CheckpointObservationPermitted next actionDecision basisUnresolved work
Before exposureR-12 passes the agreed offline and serving cases.Proceed only to the approved initial exposure stage.Evidence covers this manifest and tested scope.Live task distribution and reviewer effort remain uncertain.
Shadow comparisonCandidate outputs are recorded without changing the live route.Compare with the existing path on permitted representative data.Shadow execution must suppress consequential downstream effects.A clean shadow run does not prove users will handle the new output correctly.
Limited live useService latency is acceptable; one task category shows higher review demand.Hold expansion while investigating the category.Review burden is part of the agreed operating criteria.Outcome labels for recent cases have not yet arrived.
WithdrawalThe owner decides the observed behavior exceeds the permitted boundary.Stop R-12 exposure and use the tested previous or ordinary workflow.Recovery must account for compatibility and pending requests.Existing R-12 outputs are still present in stored records.
ReconciliationAffected records can be located by release and request identity.Route them through the agreed review or correction process.People retain authority over consequential corrections.A rollback is not complete evidence that every downstream effect was reversed.
New candidateThe cause, fix, and regression cases are documented under a new manifest.Repeat relevant acceptance before renewed exposure.A fresh release decision uses current evidence.Renaming R-12 or changing a threshold without evaluation is insufficient.

Exposure sequence

Increase use only when the next question has an answer.

Choose each stage's population, duration, evidence, and stop conditions for the workload. Low traffic, delayed outcomes, and rare failures can make a short observation window inconclusive.

  1. 01

    Freeze the candidate

    Record the complete release manifest and acceptance scope. For a hosted model, record the provider version and the limits of version pinning or reproducibility.

  2. 02

    Validate the runtime

    Exercise the actual input transformation, serving mode, security checks, output contract, capacity limits, and recovery behavior in an approved environment.

  3. 03

    Observe without effect

    Where appropriate and authorized, compare candidate outputs without applying them. Verify that duplicate traffic does not also duplicate writes, messages, or other effects.

  4. 04

    Limit live exposure

    Use a representative bounded cohort with release-specific monitoring and a tested stop path. Hold the stage when the evidence is insufficient; elapsed time alone is not acceptance.

  5. 05

    Hand over operation

    Confirm alert ownership, capacity and cost limits, model-change review, incident response, and retirement. Keep checking the task after the deployment pipeline has finished.

Production limits

Keep security, cost, and quality decisions explicit.

An operating model needs more than a dashboard. Each limit should connect an observable condition to an owner and a response.

Protect every processing path
Review inference, feature retrieval, captured samples, logs, support access, and backups. Minimize retained sensitive material and enforce authorization outside the model. Test malformed and hostile inputs alongside ordinary cases.
Budget for completed work
Measure resource use across inference, retries, retrieval, storage, monitoring, and human review. Report failed and abandoned work separately so a low cost per API call does not conceal an expensive completed task.
Treat drift as an investigation signal
A shift in inputs does not by itself quantify a loss in task quality, and stable inputs do not establish correctness. Inspect relevant outcomes and labels, their coverage, and changes in how people use the system before deciding to retrain or replace it.
Retire the whole release path
Stop scheduled jobs and traffic, resolve queued work, revoke unused access, and handle retained artifacts under the agreed policy. Preserve the evidence needed to interpret past outputs without keeping an unnecessary live dependency.

Release review questions

Resolve the assumptions hidden by a successful demo.

The answers depend on the exact workflow and serving environment.

Do all models need a real-time endpoint?
No. Scheduled batch results may fit work that tolerates delay. A queued job can provide durable status for longer processing. Interactive endpoints fit bounded response needs. Edge execution adds device resources, update delivery, and offline behavior to the contract. Choose through the workflow's timing and operating constraints.
Is the previous model enough for rollback?
Only if its inputs, runtime, configuration, and downstream contracts still work. Rehearse the actual recovery path, including pending requests and stored outputs. A previous artifact that cannot run against current data is not a usable fallback.
What if outcome labels arrive later?
State the observation delay and label coverage explicitly. Use service and input monitoring for what they can establish, while retaining a process to match later outcomes to predictions. Do not report unobserved recent cases as proven correct.
Can retraining run automatically?
Training can be automated without automatically approving deployment. A new candidate still needs appropriate data, evaluation, security, compatibility, and release checks. Decide separately who or what may promote it and under which tested conditions.
What changes when using a hosted model?
The provider operates part of the serving system, but your application still owns its input, context, output, access, and workflow behavior. Track version availability, quotas, dependency changes, and recovery options. Do not describe the full release as reproducible when an important upstream behavior cannot be pinned.

Source basis

Sources behind the control model.

  • 01

    Google Cloud

    MLOps: Continuous delivery and automation pipelines in machine learning

    Architecture guidance on the production ML lifecycle, validation, and training-serving differences. The guide does not prescribe a Google Cloud purchase.

  • 02

    Google SRE

    Canarying Releases

    Explains limited release exposure, representative observation, and release-specific evaluation. It does not establish a universal safe traffic percentage or duration.

  • 03

    Google SRE

    Data Processing Pipelines

    Supports evaluating pipeline effects and comparing outputs while suppressing production writes where the architecture permits it.

  • 04

    Amazon Web Services

    Model Monitor FAQs

    A concrete platform example of matching outcome labels to predictions and accounting for label delay. This is not a product recommendation or a claim that monitoring establishes correctness.

  • 05

    UK National Cyber Security Centre

    Guidelines for secure AI system development

    Security guidance across design, development, deployment, operation, and maintenance; implementation and testing remain necessary.

  • 06

    NIST

    AI Risk Management Framework

    Voluntary lifecycle risk-management context. The live program notes that AI RMF 1.0 is being revised; it is not evidence that this illustrative release is approved or safe.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD