Operating guide / Agent failures
Find where the work first went wrong.
An agent can produce a convincing answer while using stale evidence, changing the wrong record, or leaving work unfinished. Start with the actual outcome, reconstruct the action path, and locate the earliest supported failure. Then check which other controls allowed it to reach the user.
Begin with observable work
A bad answer is a symptom, not a diagnosis.
Separate the user's request, the evidence available at each step, the actions attempted, and the resulting state. A wrong outcome can involve several contributing failures. Changing the prompt may alter the symptom while leaving a permissive tool, stale source, or missing recovery path intact.
- Claim versus effect
- Compare the completion message with the saved artifact or system record. Requested, accepted, completed, and unknown are different states.
- First fault versus contributing fault
- A stale document may cause a wrong proposal; a missing approval check may let it become a harmful update. Investigate both.
- Containment versus repair
- Disabling an action can limit further exposure. It does not correct existing records or establish why the failure occurred.
Seven investigations
Match the repair to the failed contract.
These examples are diagnostic starting points, not a ranking of failure frequency. Preserve uncertainty when the trace cannot establish a cause. Use authorized, minimized evidence; incident logs should not become another uncontrolled copy of customer data.
Grounding: the evidence changed
01Observed symptom: The answer cites a policy, but the policy does not support the decision for this case.
- Possible mechanism
- Retrieval selected an old version, omitted a relevant exception, or lost a constraint while condensing context.
- Inspect first
- Source IDs, versions, timestamps, retrieval results, the actual context supplied, and the case facts used by the decision.
- Contain the effect
- Hold affected decisions, identify the exposed case cohort, and route uncertain cases to an owner with authoritative evidence.
- Repair owner
- The information owner and retrieval engineer establish which records apply and why.
- Durable repair
- Repair source lifecycle, filtering, context assembly, and contradiction handling. Keep factual gaps explicit instead of inviting the model to complete them.
- Misleading shortcut
- Adding more documents without checking relevance, freshness, or which passages reached the model.
- Regression proof
- Exercise expired policy, conflicting records, missing exceptions, and long-context cases; verify both cited support and the resulting decision.
Tools: success means the wrong thing
02Observed symptom: The tool returns success, but the intended update is absent or attached to another record.
- Possible mechanism
- A vague tool contract confuses submission with completion, accepts ambiguous identifiers, or hides partial results.
- Inspect first
- Tool schema and version, exact arguments, resolved record identity, raw response status, and downstream receipts.
- Contain the effect
- Suspend the affected write path and reconcile attempted operations against the destination system before retrying.
- Repair owner
- The integration owner defines operation semantics and the downstream system's acceptance evidence.
- Durable repair
- Use explicit identifiers, validated parameters, clear partial/unknown states, and independently checkable results. Make tool descriptions match actual behavior.
- Misleading shortcut
- Renaming a tool or asking the model to be careful while preserving an ambiguous API contract.
- Regression proof
- Test wrong IDs, partial batches, delayed completion, empty results, and valid alternative tool paths against actual final state.
Permissions: a proposal became authority
03Observed symptom: An action crosses the intended account, record, or approval boundary.
- Possible mechanism
- A broad service identity, trusted retrieved instruction, or reusable approval lets the agent do more than the current task permits.
- Inspect first
- Caller identity, authorization decisions, target records, approval parameters, tool access, and the untrusted content encountered.
- Contain the effect
- Disable the exposed capability through the incident process, preserve restricted evidence, and assess affected records and recipients.
- Repair owner
- Security and system owners determine containment and the required authority for each effect.
- Durable repair
- Enforce scope outside the model and bind approval to the exact action. Revalidate changed targets or parameters before execution.
- Misleading shortcut
- Treating a stronger refusal prompt, model confidence score, or valid JSON as an authorization control.
- Regression proof
- Test cross-account targets, revoked access, altered approvals, and injected instructions; denied actions must leave no unauthorized effect.
Evaluation: the test rewarded a claim
04Observed symptom: The release passed its evaluation, yet users repeatedly need to rescue nominally successful tasks.
- Possible mechanism
- The grader checks answer style, one favorable trial, or a narrow demo cohort instead of completion and residual work.
- Inspect first
- Task eligibility, test environments, grader logic, repeated trial outcomes, production exceptions, and reviewer corrections.
- Contain the effect
- Limit exposure for the affected task class and inspect a representative sample of supposedly successful work.
- Repair owner
- The product owner and evaluation lead agree what acceptable completion and escalation mean.
- Durable repair
- Add outcome checks, calibrated human review where judgment is needed, and regression cases for discovered failures. Keep failure and abandonment in reporting.
- Misleading shortcut
- Relaxing a failing assertion without proving that it rejects a valid outcome, or tuning only against the same known examples.
- Regression proof
- Show the repaired behavior across repeated and held-out cases while retaining earlier capabilities and reporting review effort.
Reliability: an uncertain effect was repeated
05Observed symptom: A timeout is followed by duplicate work, or the agent loops while the queue and cost grow.
- Possible mechanism
- Retries lack operation identity, the integration cannot resolve uncertain writes, or task limits do not stop repeated calls.
- Inspect first
- Request IDs, attempt history, receipts, timeout boundaries, queue state, latency, and the integration's retry contract.
- Contain the effect
- Pause unsafe retries, inspect destination state, and assign ambiguous operations for reconciliation.
- Repair owner
- The service owner controls recovery, retry policy, and the handling of already accepted work.
- Durable repair
- Use supported idempotency semantics, bounded attempts, explicit unknown states, and a recovery path for partial effects.
- Misleading shortcut
- Assuming a timeout means no action occurred, or assuming that rolling back application code reverses external writes.
- Regression proof
- Inject lost responses, duplicate requests, late completion, and dependency outages; inspect effects, limits, and recoverability.
Ownership: the handoff went nowhere
06Observed symptom: The agent correctly escalates, but the case remains untouched while reports count it as complete.
- Possible mechanism
- No person or queue accepts the exception, the recipient lacks evidence, or responsibility ends at tool submission.
- Inspect first
- Assignment and acknowledgement records, queue age, handoff contents, owner coverage, and the definition of completion.
- Contain the effect
- Identify stranded cases, assign an accountable owner, and correct completion reporting for the affected cohort.
- Repair owner
- The operating owner decides who accepts, resolves, and closes each exception.
- Durable repair
- Define acknowledgement, required evidence, escalation timing, and closure criteria in the workflow. Include human capacity in the operating plan.
- Misleading shortcut
- Adding another notification without establishing who must act and how acceptance is recorded.
- Regression proof
- Exercise unavailable owners, rejected handoffs, missing evidence, and reopened cases; prove each stays visible until accepted or resolved.
Change control: the tested system moved
07Observed symptom: Behavior degrades after a model, prompt, tool, policy, or data change that looked minor in isolation.
- Possible mechanism
- The release record does not identify all dependencies, or changed assumptions bypass regression and owner review.
- Inspect first
- Last accepted configuration, current component versions, source changes, release history, and outcome differences by task cohort.
- Contain the effect
- Bound exposure and compare against the last accepted system; withdraw a change only where that recovery path is valid.
- Repair owner
- The release owner coordinates engineering, information owners, and the people accepting workflow outcomes.
- Durable repair
- Record the complete tested configuration and retest when material dependencies change. Track corrective actions to closure after incidents.
- Misleading shortcut
- Blaming the latest model without checking tool, source, permission, or workflow changes made at the same time.
- Regression proof
- Replay the incident safely, run the relevant regression suite, and monitor a limited restart against explicit withdrawal conditions.
Illustrative triage notes
Keep the observation separate from the explanation.
A symptom can have several causes. These rows show what a useful incident record might contain before a team accepts a causal explanation. They are not accounts of actual incidents.
| Observed failure | Evidence to resolve | Immediate containment | Repair responsibility | Before restart |
|---|---|---|---|---|
| Wrong policy applied | Applicable version versus supplied context | Hold affected decisions | Information and retrieval owners | Exceptions and stale versions tested |
| Missing or wrong update | Target, arguments, receipt, saved state | Suspend the write path | Integration owner | Exact target and completion verified |
| Unauthorized effect | Identity, policy decision, approval scope | Disable exposed capability | Security and system owners | Negative authority tests leave no effect |
| Green eval, failed work | Grader criteria versus actual outcomes | Restrict the task cohort | Product and evaluation owners | Repeated outcome and regression checks |
| Duplicate operation | Attempts and downstream operation identity | Pause uncertain retries | Service owner | Existing effects reconciled; retry contract tested |
| Stranded escalation | Assignment, acknowledgement, queue age | Assign and track affected cases | Operating owner | Handoff acceptance and fallback exercised |
| Regression after change | Accepted and current system configuration | Reduce exposure or withdraw validly | Release owner | Causal repair and limited restart accepted |
From incident to restart
Restore control before restoring traffic.
Adapt this sequence to the organization's incident process. Urgent containment may precede complete diagnosis, but evidence and already-started work still need an owner.
- 01
Bound the exposure
Identify affected tasks, records, capabilities, and time range. Stop further harmful effects through the appropriate operational authority, while preserving the evidence needed to investigate.
- 02
Reconstruct the path
Join the request, release identity, supplied evidence, tool arguments, policy decisions, receipts, and actual outcome. Record unknowns rather than filling gaps with the agent's explanation.
- 03
Test the causal account
Reproduce safely with controlled inputs. Locate the first supported divergence and the downstream controls that failed to catch it. A correlation with the latest release is a lead, not proof.
- 04
Repair and reconcile
Fix the contract, add a regression case, and resolve existing partial or wrong effects. A software repair and a corrected customer record are separate acceptance items.
- 05
Resume with an owner
Review evidence, remaining limits, monitoring, and withdrawal conditions. Start with bounded exposure and track corrective actions until they are accepted and closed.
Evidence worth keeping
Make the next failure easier to explain.
Collect enough evidence to connect behavior to effects without retaining every sensitive input indefinitely. Set access and retention deliberately.
- A complete release identity
- Record the model and settings, prompt, tool schemas, retrieval configuration, policy version, and application release that formed the tested system.
- Action and outcome records
- Join attempts to destination receipts and verified state. Preserve partial, denied, cancelled, and unknown outcomes as distinct records.
- Service and work measures
- Watch latency, errors, load, and saturation alongside accepted work, corrections, stranded cases, and review effort. A healthy endpoint can still produce unusable work.
- Owned corrective actions
- Give each repair an owner, acceptance evidence, and review point. Share useful incident learning with appropriate data protection; avoid substituting blame for a fix.
During diagnosis
Do not let a plausible explanation end the investigation.
The strongest next step is the one that resolves uncertainty about the failure and its effects.
- Should we switch models first?
- Only when evidence points to a model limitation and a controlled comparison supports the change. A different model does not repair wrong permissions, missing records, ambiguous tool results, or an unowned queue.
- Can a prompt change be a real repair?
- Yes, if the defect is in instructions and tests establish the correction. Keep authorization, record identity, and consequential action checks in the application. A prompt cannot substitute for those controls.
- Why can a rollback leave the incident unresolved?
- Previously accepted operations, sent messages, changed records, and in-flight work can remain after code is withdrawn. Reconcile those effects separately and use the destination system's supported recovery process.
- What if we cannot reproduce the failure exactly?
- Preserve that limitation. Compare multiple trials, inspect available state and traces, and test the suspected mechanism under controlled conditions. Do not promote a hypothesis to a confirmed cause merely because it sounds convincing.
- When is it reasonable to resume?
- When the accountable owner accepts containment, repair evidence, reconciliation status, remaining uncertainty, and the bounded restart plan. A passing test is necessary evidence for the behavior it covers, not blanket proof that every task is safe.
Source basis
Sources behind the control model.
- 01
Anthropic
Effective context engineering for AI agentsContext selection and maintenance guidance. The policy-version scenario is Werkon's illustrative diagnostic analysis.
- 02
Anthropic
Writing effective tools for agentsTool purpose, descriptions, returned context, and evaluation guidance; no vendor performance result is claimed here.
- 03
Anthropic
Demystifying evals for AI agentsDistinguishes trials, traces, graders, and environmental outcomes; supports repeated and regression evaluation.
- 04
OWASP
AI Agent Security Cheat SheetCommunity control guidance for scoped tools, untrusted content, exact approvals, limits, and adversarial tests; not implementation certification.
- 05
Amazon Builders' Library
Making retries safe with idempotent APIsExplains ambiguous responses, operation identity, and retry semantics. Actual integrations must establish their own supported contract.
- 06
Google SRE
Monitoring Distributed SystemsFoundational service monitoring guidance; business outcome and handoff measures here are proposed operating checks.
- 07
Google SRE
Postmortem Culture: Learning from FailureSupports evidence-based incident review, constructive learning, protected data, and reviewed corrective actions.
Start with one real workflow
A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.
Show Us the WorkflowStart with the free automation readiness checklistOBSERVEQUANTIFYDECIDEBUILD
