Skip to main content

Exception handling agent

A detected exception is not an owned recovery.

An exception handling agent organizes unexpected operational events into evidence-based cases, finds the approved response playbook and routes work to an accountable owner. It can summarize timelines and missing context while software controls permitted actions and retries. Werkon would validate this pattern through verified recovery, with consequential decisions and safety, security or emergency response owned by qualified people.

Decide what the signal actually represents

A validation failure, expected variation, duplicate alert, maintenance condition, complaint, security incident and emergency require different handling. Preserve original messages, API problems, measurements and reports with source identity, timestamps, units and corrections. Check clock drift, stale observations and sequence before correlating events through explicit keys and time windows.

Priority candidates should show the affected work, people, money, data and time exposure. Model confidence cannot determine urgency or establish a root cause. Keep current authoritative state, observations, conflicts and unknowns separate, and retrieve only the context the case permits.

Make the playbook and handoff explicit

An approved playbook defines prerequisites, allowed inputs, reads and writes, action limits, approvals, retries, compensation, stop conditions, escalation and expiry. Start with read-only steps where possible. Only reversible, idempotent and pre-authorized low-impact actions are candidates for automatic execution.

Assign one current accountable owner and record accepted or rejected transfers, the next action, response target and fallback. Safety, medical, security and emergency signals need independent approved routes. Exception details must not expose stack traces, credentials, personal data or sensitive controls.

Reconcile recovery before closure

Read current state immediately before action and block changed preconditions or competing writes. Cap retries and reconcile an uncertain result before repeating it. Preserve manual correction, abort, rollback and safe degraded operation. Qualified owners retain continuity activation, customer remedies, compensation and final closure.

Technical success, restored service, resolved impact and confirmed root cause are separate findings. Verify recovery in the affected system and human or physical process, retain residual risk and observe recurrence. A contributing condition or hypothesis becomes a confirmed cause only through qualified investigation.

Exception boundary

Turn abnormal signals into owned work without letting the alert decide the truth.

An exception crosses system state, human judgment and physical operations. Four boundaries keep detection useful and recovery accountable.

01

Detection and source validation

Bind expected state and exception condition to an authoritative process version, then validate source, schema, units, clocks, freshness and integrity before creating or updating a case.

Required evidence: Tenant and process, work item, expected and observed states, source and event identifiers, source and receipt times, unit and schema, freshness, threshold or rule version, validation result, duplicate key, conflicting evidence and rejected signal.

02

Context, correlation, and classification

Gather only authorized source-linked context, distinguish facts from unknowns and hypotheses, correlate by domain keys and produce reviewable type, severity and routing candidates.

Required evidence: Original records, access purpose, affected people, assets, customers and work, related-case rationale, type and severity candidates, safety and security flags, material conflicts, uncertainty, missing context, correction and human classification.

03

Versioned playbook and bounded action

Match an approved current playbook, verify prerequisites and permissions immediately before each step, keep actions reversible and idempotent, and stop on changed state or uncertain effect.

Required evidence: Playbook identifier and version, scope, prerequisites, permitted tools and fields, approval, action request, idempotency key, attempt, technical response, uncertain result, retry cap, stop reason, rollback or compensation and authoritative post-action state.

04

Owned handoff, recovery, and learning

Preserve one accountable owner and every transfer, keep specialist and emergency routes independent, reconcile affected service and retain recurrence, impact, root-cause and improvement evidence beyond case closure.

Required evidence: Owner and supporting roles, route, delivery, acceptance or rejection, fallback and escalation, decision and action, service restoration, customer or physical confirmation, residual risk, incident, claim, complaint, loss, recurrence, root-cause status and verified change.

Signal-to-outcome path

Keep detection, case ownership, action, and recovery as separate states.

A signal can start investigation. It cannot grant authority, prove impact or close the affected work.

  1. 01

    Resolve expected state and evidence

    Identify tenant, process, work item and authoritative state; capture original observations; validate source, schema, units, sequence and freshness; and hold invalid or conflicting signals for review.

    Owner
    Process, system, data and monitoring owners
    Evidence
    Expected state, state source and version, observation, event or problem record, timestamps, units, integrity and access context, validation, conflict, expiry and correction.
  2. 02

    Correlate and open accountable work

    Apply deterministic domain keys and bounded windows, preserve related candidates, create one attributable case and produce type, priority and specialist-route candidates without inventing urgency.

    Owner
    Operations, service, safety and security triage owners
    Evidence
    Correlation keys and rationale, duplicate or related state, case identifier, affected scope, facts, unknowns, type and severity candidate, safety or security flag, owner route and manual correction.
  3. 03

    Select the bounded playbook

    Find an approved current playbook, check type and scope, permissions, prerequisites, dependencies, stop conditions, approval points, retry and compensation rules, and fall back to human handling when any contract is missing.

    Owner
    Process, control, safety, security and change owners
    Evidence
    Playbook and version, approval and expiry, matched scope, required state, permissions, allowed tools and writes, stop rules, retry cap, compensation, escalation and selection decision.
  4. 04

    Act, observe, and hand off

    Re-read current state, perform only pre-authorized reversible steps, record each attempt and uncertain result, stop rather than retry blindly, and transfer unresolved work to an accepting named owner.

    Owner
    Operations owner and qualified specialist responders
    Evidence
    Pre-action state, authority, action request, idempotency key, attempt and response, changed state, conflict, stop or rollback, handoff packet, delivery, acceptance, rejection, fallback and next action.
  5. 05

    Reconcile recovery and recurrence

    Verify affected system and physical or human service, keep residual impact open, require qualified root-cause work and compare recurrence, customer impact and human effort before expanding automation.

    Owner
    Service, continuity, customer, risk and improvement owners
    Evidence
    Authoritative recovered state, physical or customer confirmation, unresolved impact, incident, claim, complaint, loss, recurrence, contributing conditions, root-cause status, corrective action, review and verified outcome.

Authority map

Separate exception assistance, deterministic controls, and accountable response.

A model can summarize an abnormal trace. It cannot declare an incident, authorize a dangerous action or decide the business has recovered.

01

Deterministic exception controls

Software owns tenant boundaries, source and schema validation, clocks, state authority, correlation keys, thresholds, hard routes, permissions, playbook versions, prerequisites, idempotency, retry caps, case transitions, deadlines and receipts.

  • Tenant, process, work, source, event, case and playbook identifiers
  • Freshness, units, sequence, duplicate, correlation, state and threshold validation
  • Permissions, prerequisites, allowed writes, idempotency, rate, retry, stop and rollback gates
  • Detection, case, owner, action, handoff, recovery and closure receipts
02

Bounded AI assistance

Models can classify a candidate, extract evidence with source references, summarize a timeline, identify missing context, find an approved playbook and draft clarification or handoff text while truth, priority, routing and action authority stay outside the model.

  • Exception-type and specialist-route candidates
  • Source-linked entities, facts, conflicts, unknowns and impact candidates
  • Related-case and approved-playbook retrieval candidates
  • Timeline, clarification, handoff and review drafts
03

Accountable operational authority

Named operations owners and qualified safety, security, continuity, compliance, customer, finance, legal and people responders own material classification, priority, action, emergency response, customer remedy, root cause, residual risk and closure.

  • Severity, incident, safety, security and continuity decisions
  • Consequential system, financial, contractual, customer and production action
  • Exception acceptance, investigation, escalation, recovery and residual-risk decisions
  • Root cause, corrective action, claim, complaint, compensation and final closure

Exception-handling components

Build an evidence and recovery ledger, not an alert inbox.

Reliable handling depends on exact state, controlled action and durable ownership. Four components preserve that chain.

01

Expected-state and observation ledger

Version tenant, process, work item, expected state, authoritative source, exception condition, original observation, source and receipt times, schema, units, freshness, integrity, validation and correction.

Operating contract: Event is not business state, metric is not impact, threshold breach is not incident, model score is not fact, receipt time is not occurrence time, silence is not success and a monitoring gap is not normal operation.

02

Case and correlation ledger

Preserve case type, domain keys, correlation window and rationale, duplicates and related cases, affected scope, facts, conflicts, unknowns, severity candidate, hard route, current owner, status and response policy.

Operating contract: Similarity is not identity, shared timing is not shared cause, duplicate is not harmless, many alerts are not many exceptions, priority candidate is not authority, routed is not delivered, delivered is not accepted and aging is not owner action.

03

Playbook and action ledger

Bind approved playbook version, prerequisites, permissions, tools, reads and writes, approval points, action request, idempotency, attempt, result, uncertainty, retry cap, stop, rollback and compensation to current source state.

Operating contract: Runbook text is not current approval, technical response is not business success, timeout is not failure, failure is not safe to retry, repeated request is not idempotent by intent alone, rollback request is not restored state and automation rate is not value.

04

Handoff, recovery, and learning ledger

Link owner transfer, acceptance, specialist or emergency route, decision, affected service, physical or customer confirmation, residual impact, incident, claim, complaint, recurrence, root-cause status and corrective action.

Operating contract: Case closure is not service recovery, restoration is not root cause, no recurrence yet is not prevention, compensation is not absence of harm, completed corrective task is not verified improvement and continuity invocation is not proof of resilience.

Delivery path

Prove one exception family and playbook before expanding.

Begin where expected state, owners, actions and recovery can be inspected without placing safety or continuity decisions inside the agent.

  1. 01

    Choose one bounded exception family

    Select one tenant, process, work-item class, authoritative source, exception condition, owner group and low-impact response with enough historical evidence to inspect missed, false and repeated cases.

  2. 02

    Map states, evidence, and authority

    Inventory expected and observed states, schemas, units, clocks, freshness, duplicates, correlations, priorities, hard specialist routes, owner transitions, response policies, permissions, approvals and recovery evidence.

  3. 03

    Encode versioned bounded playbooks

    Implement source validation, deterministic correlation and case state, minimal context retrieval, explicit prerequisites, pre-authorized reversible actions, idempotency, retry caps, stops, compensation, handoffs and receipts.

  4. 04

    Test failure inside the handler

    Exercise forged and stale sources, duplicates, clock drift, conflicting state, unsafe merges, stale playbooks, revoked permission, partial writes, timeout after commit, concurrent responders, unavailable owners, emergency signals and failed rollback.

  5. 05

    Release read-first and reconcile outcomes

    Start with detection and owner-visible drafts, compare corrections, accepted cases, safe actions, unresolved age, recovery, recurrence, customer impact and human effort, then expand only steps with current evidence and reliable stop paths.

Release controls

Six controls before an exception handler can change operations.

Fast action can deepen a bad signal. These controls keep detection, playbooks and ownership honest.

Expected state and source are explicit
Bind every condition to tenant, process, work item, authoritative state and policy version; preserve original observation, schema, units and clocks; reject stale or impossible input; and keep conflicting evidence visible.
Correlation is deterministic and reversible
Use durable domain keys and bounded windows, record why records were joined, retain related candidates, prevent cross-tenant merges and support safe split, correction and deduplication without losing source evidence.
Priority and specialist routes stay governed
Use explicit impact dimensions and hard safety, emergency, security and continuity routes; present severity only as a candidate; keep material declaration, customer commitment and consequential response with qualified owners.
Playbooks are approved and state-aware
Version scope, prerequisites, permissions, reads, writes, approvals, expiry, stop, retry and compensation; re-read authoritative state before action; block stale or mismatched playbooks; and never let generated text expand authority.
Actions are bounded, attributable, and recoverable
Prefer read-only work, limit automatic steps to reversible low-impact actions, use idempotency and rate controls, preserve attempts and uncertain results, reconcile before retry and provide independent abort, rollback and human override.
Ownership and recovery remain observable
Require one current owner, acceptance and deterministic fallback; distinguish action from recovery; verify affected service and physical work; retain residual impact, recurrence, claims and complaints; and prohibit ownerless autonomous closure.

Outcome evidence

Measure owned recovery and prevented harm, not alerts closed.

Closing more cases can hide missed exceptions, unsafe merges and false recovery. Proof must follow the affected operation.

Baseline

  • Processes, work-item classes, expected-state sources, signal types, thresholds, schemas, freshness windows, exception families, owners, specialist routes, playbooks, permissions, action classes and recovery definitions
  • Current operator and specialist time from signal review through validation, correlation, case acceptance, investigation, action, handoff, service reconciliation, root-cause work and improvement review
  • Current missed and false detections, duplicates, unsafe merges, priority corrections, stale playbooks, blocked and uncertain actions, retries, owner rejections, handoff delays, unresolved aging, recovery failures and recurrence
  • Current affected service, people, assets, customers, money and data plus incidents, claims, complaints, rework, loss, degraded operation, recovery, residual risk and verified corrective changes

Outcome evidence

  • Correct source, exception type, correlation, severity candidate, playbook, permission, owner, action and recovery state against authoritative evidence
  • Valid detection, duplicate prevention, owner acceptance, safe bounded action, timely specialist route, reconciled recovery and verified corrective change by exception family and operating context
  • Missed harmful exception, false case, cross-tenant or unsafe merge, privacy leak, stale action, blind retry, lost handoff, ownerless case and false recovery prevention
  • Operator effort, queue displacement, degraded time, service restoration, recurrence, customer impact, incidents, claims, complaints, loss and residual risk against the prior process with demand and source quality visible

Guardrails

  • Wrong tenant or work item, forged or stale source, invalid units, clock error, exposed sensitive detail, cross-case leakage, unsafe merge, hidden conflict, inaccessible correction and missing original evidence
  • Model score treated as fact or urgency, threshold called incident, alert volume called impact, generated root cause, ignored hard route, unavailable owner, unaccepted handoff and safety, security or continuity response delayed for AI
  • Stale or unapproved playbook, changed prerequisite, excess permission, irreversible automatic action, duplicate write, timeout retried blindly, uncertain effect hidden, rollback unavailable and handler failure not opened as its own exception
  • Technical success called recovery, closed ticket called resolved impact, restoration called root cause, recurring case suppressed, customer or physical state unreconciled and automation rate presented as saving, resilience or service improvement

Fit test

Use this pattern when state authority, playbooks, and owners are inspectable.

Good reason to begin

  • One operation has durable tenant, process and work-item identifiers, explicit expected state, authoritative current-state sources, observable exceptions and enough history to inspect corrections and recurrence.
  • Exception types, impact dimensions, hard specialist routes, current owners, acceptance, fallback, recovery definitions and case transitions are explicit and maintained.
  • A narrow playbook is approved, versioned, permission-scoped, state-aware, reversible, idempotent and testable with independent stop, rollback and human override.
  • Affected service and physical or customer outcomes can be reconciled beyond the ticket, while incident, claim, complaint, residual risk, root cause and corrective action remain attributable.

Resolve before beginning

  • Expected state, source authority, timestamps, units, exception condition, correlation keys, sensitive-data scope, safety or security routes, owner or recovery evidence is undefined.
  • The process cannot distinguish observation from exception, duplicate from related case, severity candidate from declaration, action attempt from effect, restoration from root cause or case closure from recovery.
  • Success is defined by alert reduction, automation rate or closure speed without missed harm, false cases, unsafe merges, operator burden, unresolved age, recurrence, customer impact and physical service evidence.
  • The agent is expected to diagnose incidents, choose material priority, activate emergency or continuity response, expose internals, run irreversible actions, retry uncertain writes, resolve ownerless work or fabricate recovery.

Source basis

Sources behind the control model.

  • 01

    Object Management Group

    Business Process Model and Notation, version 2.0.2

    OMG lists BPMN 2.0.2 as its latest formal version and provides the normative specification. BPMN models error events that interrupt an activity and escalation events that can interrupt or continue alongside it. A process notation does not prove a runtime signal, current process deployment, exception type, priority, owner acceptance, safe action, recovery or outcome.

  • 02

    RFC Editor

    RFC 9457: Problem Details for HTTP APIs

    This Internet Standards Track document, which obsoletes RFC 7807, defines machine-readable problem details for HTTP APIs. It says the status member is advisory and warns that problem details can expose system, access and privacy information. An HTTP problem does not establish domain state, root cause, business severity, case ownership, physical effect or recovery.

  • 03

    Cloud Native Computing Foundation

    CloudEvents 1.0.2

    The CloudEvents project identifies 1.0.2 as the latest released core specification for describing event data in a common way. A portable envelope does not prove source identity, payload correctness, semantic uniqueness, ordering, complete coverage, delivery, exactly-once processing, current business state, exception priority, owner action or outcome.

  • 04

    International Organization for Standardization

    ISO 22301:2019: Business continuity management systems

    ISO lists this business-continuity management-system requirements standard as published while marked to be revised. Its public description covers planning, operation, monitoring, review, maintenance and improvement for disruptive incidents and recovery. It does not detect a specific exception, validate a signal, define a playbook, authorize continuity activation, prove recovery or certify this agent.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD