Skip to main content

Engineering guide / business agents

Build the execution contract before the agent loop.

A business agent needs more than a model that selects tools. It needs a trustworthy account of what was requested, what was permitted, what changed, and what remains unresolved. Start with one task whose outcome you can verify, then design the loop around that contract.

Start with the task

Give the model choices inside a process you can account for.

An agent uses model output to choose its next step from available actions, observes the result, and continues within limits. A fixed workflow follows a path defined in code. Use the agent pattern when choosing the next step adds value, and keep predictable validation and execution in ordinary software.

A workflow may be enough
If every ticket requires the same lookup, template, and review, implement that sequence directly. Dynamic tool selection earns its place only if representative tasks demonstrate a useful improvement over the simpler baseline.
The loop is only one component
Context preparation, model calls, tool adapters, state storage, review screens, and operational controls all affect the result. A framework can connect them, but its name does not establish their behavior.
Completion is a checked condition
For our example, success means one approved internal note exists on the intended ticket with the expected content. A model saying it finished, or an API accepting a job, is insufficient evidence of that condition.

Six implementation boundaries

Build the parts that make a run explainable.

Consider an internal support assistant that reads an authorized ticket and drafts a handover note. A reviewer may approve adding that note. Sending a customer message, changing priority, assigning work, or closing the ticket are outside this example. The following artifacts describe a proposed design, not a ready-made integration.

Request and run identity

01

Design question: Which task did this person actually authorize?

Example contract
Read one selected ticket and prepare an internal note. Adding the note requires a separate approval.
Inputs
Authenticated actor, tenant, selected ticket, requested outcome, and task limits.
Implementation
Create a run record in trusted application code. Derive identity from the session, not model-supplied fields. Attach a deadline and step budget.
Owner
The product owner defines the supported task; system policy controls access.
Acceptance test
A request naming a ticket from another tenant is rejected before its content enters model context.
Failure to expose
A plausible identifier is treated as proof of access.
Retained artifact
A run identifier bound to actor, tenant, task, and policy version.

Context with source boundaries

02

Design question: What material can inform this draft?

Example contract
Only permitted ticket events and the approved handover format enter the working context.
Inputs
Source identifiers, versions, relevant text, and retrieval access checks.
Implementation
Keep retrieved instructions as untrusted content. Store task state separately from generated summaries; do not let remembered text grant permissions.
Owner
Data owners define access and retention; reviewers resolve missing or conflicting facts.
Acceptance test
A ticket containing instructions to export other records cannot expand the tool scope.
Failure to expose
A summary loses a qualification and later appears to be an authoritative fact.
Retained artifact
A bounded context manifest linking draft statements to source events.

A bounded decision loop

03

Design question: What can the model choose next?

Example contract
Read an allowed event, prepare a draft, request clarification, or stop.
Inputs
Task contract, available tool descriptions, prior observations, remaining budget.
Implementation
Validate each proposed call before dispatch. Count retries and model calls against run limits. Return explicit observations so the loop can distinguish missing data from tool failure.
Owner
The runtime enforces budgets and permitted transitions, including a blocked outcome.
Acceptance test
Repeated identical lookups exhaust the configured limit and produce an actionable handoff.
Failure to expose
A loop continues spending resources because its final answer has not arrived.
Retained artifact
A sequence of selected actions and observations without requiring hidden model reasoning.

Tools with narrow effects

04

Design question: What exactly can this adapter change?

Example contract
The note tool can append an internal note to the selected ticket; it cannot send messages or alter unrelated fields.
Inputs
Validated ticket identifier, proposed note, source version, and approval reference.
Implementation
Use a typed input schema and explicit error outcomes. Enforce resource authorization at execution. Keep credentials in the service boundary and exclude them from model-visible results.
Owner
The integration owner defines tool semantics and downstream permissions.
Acceptance test
Extra fields, unauthorized records, and expired access are rejected without a write.
Failure to expose
A general-purpose update function allows more effects than its label suggests.
Retained artifact
Versioned tool contracts and tests for allowed and denied effects.

Approval and durable execution

05

Design question: Can an interrupted run establish what happened?

Example contract
Approval binds the exact note, destination, and relevant source version. Changed content requires a new decision.
Inputs
Proposal snapshot, reviewer identity, validity period, and operation identity.
Implementation
Persist execution intent before dispatch. Record the downstream receipt when available. A timeout becomes an unknown outcome requiring reconciliation; do not silently start a fresh write.
Owner
The reviewer decides the proposed change; the runtime checks that the approval still applies.
Acceptance test
Interrupt the worker after the write but before receipt storage, then recover without appending a duplicate.
Failure to expose
A second attempt creates another note while the first note already exists.
Retained artifact
An action record separating proposed, approved, in-flight, unknown, and confirmed states.

Task evaluation and operation

06

Design question: Does the system finish the right work, including exceptions?

Example contract
Assess note fidelity, scope compliance, verified writes, reviewer corrections, and unresolved runs separately.
Inputs
Representative tickets, expected outcomes, adversarial cases, and controlled fault scenarios.
Implementation
Inspect both execution traces and final system state. Repeat cases to expose variability. Keep a held-out set when tuning prompts or tools.
Owner
A named release owner decides whether evidence supports the permitted operating scope.
Acceptance test
A fluent summary cannot pass a case where the note was posted to the wrong ticket.
Failure to expose
A text-quality score hides a failed or unauthorized downstream action.
Retained artifact
Versioned evaluation results, known limits, an incident path, and an operating owner.

Worked execution trace

A timeout is not proof that the write failed.

In this illustrative run, ticket T-184 is at version 7 and a reviewer approves note proposal P-3. The downstream API must support the proposed operation identity and recovery contract. If it cannot establish whether a write occurred, keep the run unresolved for an operator instead of assuming success or retrying blindly.

StateRecorded evidenceAllowed next stepEnforced conditionWhat prevents progress
Draft preparedP-3 contains the exact note, ticket T-184, and source version 7.Present the draft and its supporting events.No write authority exists yet.A missing event or unsupported statement needs correction.
Approval recordedReviewer accepts P-3 for the stated destination and effect.Request dispatch of the approved proposal.Check current permission, approval validity, and relevant record version.Version 8 or changed note content invalidates this proposal's execution path.
Write dispatchedPersisted operation A-9 binds this effect before the call.Wait for the adapter observation.Only the approved payload is sent.A timeout leaves the outcome unknown.
Outcome unknownA-9 exists locally; no conclusive receipt is stored.Use the documented status lookup or hand off.Recovery follows the API's deduplication and retention contract.No reliable lookup or expired deduplication coverage requires operator resolution.
Effect confirmedAuthoritative lookup links A-9 to note N-52 on T-184 with matching content.Report the confirmed note and its reference.Completion is derived from checked evidence.Wrong content or destination is an incident, not a successful run.
Run closedThe requested note exists; no additional action was authorized.Show the final outcome and any remaining limitations.Do not infer that the ticket is resolved or a customer was contacted.Any further action requires its own task contract.

Development sequence

Test recovery before expanding authority.

Build one complete path through the real boundaries before adding more tools. This sequence is an engineering proposal; the release decision depends on the workflow's consequences and measured evidence.

  1. 01

    Specify the outcome

    Write success, blocked, rejected, and unknown examples. Establish a manual or fixed-workflow baseline and identify what model-directed choices should improve.

  2. 02

    Build a controlled adapter

    Implement contracts against a test system. Exercise denied access, invalid payloads, duplicate calls, concurrent changes, and ambiguous responses before wiring the model loop.

  3. 03

    Evaluate whole runs

    Test ordinary and difficult tasks, instruction attacks, stale approvals, worker restarts, and partial failures. Inspect the final record as well as the conversation.

  4. 04

    Introduce supervised use

    Begin with read and draft behavior under an approved data boundary. Measure corrections and unresolved work before considering a narrowly approved write path.

  5. 05

    Operate a versioned release

    Track model, prompt, context, tool, and policy versions. Re-evaluate material changes, rehearse interruption and recovery, and retain a route back to the ordinary workflow.

Decisions outside the loop

Make the operational exceptions visible.

These rules connect the example to everyday support work. They need explicit product behavior, not just instructions in a system prompt.

Approval is attached to an effect
Show the destination, full note, source context, and unresolved issues together. Rejecting a proposal must end that path. An edited proposal returns to review rather than inheriting an earlier approval.
Concurrency is part of correctness
A read followed by a write can race with another worker. Where supported, use a server-enforced conditional write such as If-Match, or an equivalent transaction/version check. A client-side comparison alone does not close that gap.
Cancellation has a boundary
Stopping a run can prevent future calls but may not undo a dispatched write. Show any in-flight operation and reconcile it. A compensating action has its own permissions and consequences; it is not an automatic rollback.
Evidence has its own access rules
Operators need event identifiers, policy decisions, proposal versions, and receipts. They do not need every sensitive source copied into every log. Define access, retention, redaction, and incident evidence needs for each record.

Implementation questions

Resolve these before choosing the runtime.

The useful comparison is whether your team can implement, inspect, test, and operate the required behavior.

Do we need an agent framework?
Not necessarily. A small explicit loop can be easier to inspect for one bounded task. A framework may help with persistence, tool orchestration, or tracing, but verify its retry, cancellation, authorization, and versioning behavior against your own contracts.
Should the first version use several agents?
Only if a measured task need justifies the extra coordination. Start with one accountable run and a limited tool set. Additional agents introduce more handoffs, context boundaries, and failure states; agreement between models does not establish authority.
Does a tool schema make a call safe?
No. It can establish shape, not permission, factual correctness, or the acceptability of an effect. A correctly formed call can still target the wrong record. Validate semantics and authorization in trusted code at the execution boundary.
What if the system has no idempotent write API?
Do not promise duplicate-free retries by adding a local key alone. Determine whether the downstream system can recognize an operation or expose conclusive status. Where ambiguity cannot be resolved, retain an operator-controlled path and keep the outcome explicitly unknown.
What should a release report contain?
The tested task scope, dataset and version details, success and failure definitions, observed failures, reviewer workload, resource use, recovery results, owners, and unresolved limits. A single aggregate score is too little to decide which actions may run.

Source basis

Sources behind the control model.

  • 01

    Anthropic

    Building effective agents

    Vendor engineering perspective on workflows, agents, simple architectures, and tool interfaces. The article notes that its tooling landscape has changed since 2024.

  • 02

    Anthropic

    Demystifying evals for AI agents

    Supports evaluating trajectories and outcomes across repeated tasks rather than judging final prose alone.

  • 03

    OWASP

    AI Agent Security Cheat Sheet

    Security guidance for tool scope, untrusted context, memory, approvals, and abuse-case testing. Guidance is not evidence that a specific implementation is secure.

  • 04

    OWASP

    Authorization Cheat Sheet

    Basis for enforcing least privilege and checking access at the application boundary, independently of model output.

  • 05

    Amazon Builders' Library

    Making retries safe with idempotent APIs

    Explains ambiguous outcomes, request identity, duplicate effects, and the importance of the service's idempotency contract. The ticket trace is Werkon's illustrative application.

  • 06

    IETF

    RFC 9110: HTTP Semantics

    Defines If-Match and conditional request semantics. An integration must actually support and enforce the relevant precondition.

  • 07

    UK National Cyber Security Centre

    Guidelines for secure AI system development

    Lifecycle guidance spanning design, development, deployment, and operation. It does not certify the architecture described here.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD