Engineering guide / business agents
Build the execution contract before the agent loop.
A business agent needs more than a model that selects tools. It needs a trustworthy account of what was requested, what was permitted, what changed, and what remains unresolved. Start with one task whose outcome you can verify, then design the loop around that contract.
Start with the task
Give the model choices inside a process you can account for.
An agent uses model output to choose its next step from available actions, observes the result, and continues within limits. A fixed workflow follows a path defined in code. Use the agent pattern when choosing the next step adds value, and keep predictable validation and execution in ordinary software.
- A workflow may be enough
- If every ticket requires the same lookup, template, and review, implement that sequence directly. Dynamic tool selection earns its place only if representative tasks demonstrate a useful improvement over the simpler baseline.
- The loop is only one component
- Context preparation, model calls, tool adapters, state storage, review screens, and operational controls all affect the result. A framework can connect them, but its name does not establish their behavior.
- Completion is a checked condition
- For our example, success means one approved internal note exists on the intended ticket with the expected content. A model saying it finished, or an API accepting a job, is insufficient evidence of that condition.
Six implementation boundaries
Build the parts that make a run explainable.
Consider an internal support assistant that reads an authorized ticket and drafts a handover note. A reviewer may approve adding that note. Sending a customer message, changing priority, assigning work, or closing the ticket are outside this example. The following artifacts describe a proposed design, not a ready-made integration.
Request and run identity
01Design question: Which task did this person actually authorize?
- Example contract
- Read one selected ticket and prepare an internal note. Adding the note requires a separate approval.
- Inputs
- Authenticated actor, tenant, selected ticket, requested outcome, and task limits.
- Implementation
- Create a run record in trusted application code. Derive identity from the session, not model-supplied fields. Attach a deadline and step budget.
- Owner
- The product owner defines the supported task; system policy controls access.
- Acceptance test
- A request naming a ticket from another tenant is rejected before its content enters model context.
- Failure to expose
- A plausible identifier is treated as proof of access.
- Retained artifact
- A run identifier bound to actor, tenant, task, and policy version.
Context with source boundaries
02Design question: What material can inform this draft?
- Example contract
- Only permitted ticket events and the approved handover format enter the working context.
- Inputs
- Source identifiers, versions, relevant text, and retrieval access checks.
- Implementation
- Keep retrieved instructions as untrusted content. Store task state separately from generated summaries; do not let remembered text grant permissions.
- Owner
- Data owners define access and retention; reviewers resolve missing or conflicting facts.
- Acceptance test
- A ticket containing instructions to export other records cannot expand the tool scope.
- Failure to expose
- A summary loses a qualification and later appears to be an authoritative fact.
- Retained artifact
- A bounded context manifest linking draft statements to source events.
A bounded decision loop
03Design question: What can the model choose next?
- Example contract
- Read an allowed event, prepare a draft, request clarification, or stop.
- Inputs
- Task contract, available tool descriptions, prior observations, remaining budget.
- Implementation
- Validate each proposed call before dispatch. Count retries and model calls against run limits. Return explicit observations so the loop can distinguish missing data from tool failure.
- Owner
- The runtime enforces budgets and permitted transitions, including a blocked outcome.
- Acceptance test
- Repeated identical lookups exhaust the configured limit and produce an actionable handoff.
- Failure to expose
- A loop continues spending resources because its final answer has not arrived.
- Retained artifact
- A sequence of selected actions and observations without requiring hidden model reasoning.
Tools with narrow effects
04Design question: What exactly can this adapter change?
- Example contract
- The note tool can append an internal note to the selected ticket; it cannot send messages or alter unrelated fields.
- Inputs
- Validated ticket identifier, proposed note, source version, and approval reference.
- Implementation
- Use a typed input schema and explicit error outcomes. Enforce resource authorization at execution. Keep credentials in the service boundary and exclude them from model-visible results.
- Owner
- The integration owner defines tool semantics and downstream permissions.
- Acceptance test
- Extra fields, unauthorized records, and expired access are rejected without a write.
- Failure to expose
- A general-purpose update function allows more effects than its label suggests.
- Retained artifact
- Versioned tool contracts and tests for allowed and denied effects.
Approval and durable execution
05Design question: Can an interrupted run establish what happened?
- Example contract
- Approval binds the exact note, destination, and relevant source version. Changed content requires a new decision.
- Inputs
- Proposal snapshot, reviewer identity, validity period, and operation identity.
- Implementation
- Persist execution intent before dispatch. Record the downstream receipt when available. A timeout becomes an unknown outcome requiring reconciliation; do not silently start a fresh write.
- Owner
- The reviewer decides the proposed change; the runtime checks that the approval still applies.
- Acceptance test
- Interrupt the worker after the write but before receipt storage, then recover without appending a duplicate.
- Failure to expose
- A second attempt creates another note while the first note already exists.
- Retained artifact
- An action record separating proposed, approved, in-flight, unknown, and confirmed states.
Task evaluation and operation
06Design question: Does the system finish the right work, including exceptions?
- Example contract
- Assess note fidelity, scope compliance, verified writes, reviewer corrections, and unresolved runs separately.
- Inputs
- Representative tickets, expected outcomes, adversarial cases, and controlled fault scenarios.
- Implementation
- Inspect both execution traces and final system state. Repeat cases to expose variability. Keep a held-out set when tuning prompts or tools.
- Owner
- A named release owner decides whether evidence supports the permitted operating scope.
- Acceptance test
- A fluent summary cannot pass a case where the note was posted to the wrong ticket.
- Failure to expose
- A text-quality score hides a failed or unauthorized downstream action.
- Retained artifact
- Versioned evaluation results, known limits, an incident path, and an operating owner.
Worked execution trace
A timeout is not proof that the write failed.
In this illustrative run, ticket T-184 is at version 7 and a reviewer approves note proposal P-3. The downstream API must support the proposed operation identity and recovery contract. If it cannot establish whether a write occurred, keep the run unresolved for an operator instead of assuming success or retrying blindly.
| State | Recorded evidence | Allowed next step | Enforced condition | What prevents progress |
|---|---|---|---|---|
| Draft prepared | P-3 contains the exact note, ticket T-184, and source version 7. | Present the draft and its supporting events. | No write authority exists yet. | A missing event or unsupported statement needs correction. |
| Approval recorded | Reviewer accepts P-3 for the stated destination and effect. | Request dispatch of the approved proposal. | Check current permission, approval validity, and relevant record version. | Version 8 or changed note content invalidates this proposal's execution path. |
| Write dispatched | Persisted operation A-9 binds this effect before the call. | Wait for the adapter observation. | Only the approved payload is sent. | A timeout leaves the outcome unknown. |
| Outcome unknown | A-9 exists locally; no conclusive receipt is stored. | Use the documented status lookup or hand off. | Recovery follows the API's deduplication and retention contract. | No reliable lookup or expired deduplication coverage requires operator resolution. |
| Effect confirmed | Authoritative lookup links A-9 to note N-52 on T-184 with matching content. | Report the confirmed note and its reference. | Completion is derived from checked evidence. | Wrong content or destination is an incident, not a successful run. |
| Run closed | The requested note exists; no additional action was authorized. | Show the final outcome and any remaining limitations. | Do not infer that the ticket is resolved or a customer was contacted. | Any further action requires its own task contract. |
Development sequence
Test recovery before expanding authority.
Build one complete path through the real boundaries before adding more tools. This sequence is an engineering proposal; the release decision depends on the workflow's consequences and measured evidence.
- 01
Specify the outcome
Write success, blocked, rejected, and unknown examples. Establish a manual or fixed-workflow baseline and identify what model-directed choices should improve.
- 02
Build a controlled adapter
Implement contracts against a test system. Exercise denied access, invalid payloads, duplicate calls, concurrent changes, and ambiguous responses before wiring the model loop.
- 03
Evaluate whole runs
Test ordinary and difficult tasks, instruction attacks, stale approvals, worker restarts, and partial failures. Inspect the final record as well as the conversation.
- 04
Introduce supervised use
Begin with read and draft behavior under an approved data boundary. Measure corrections and unresolved work before considering a narrowly approved write path.
- 05
Operate a versioned release
Track model, prompt, context, tool, and policy versions. Re-evaluate material changes, rehearse interruption and recovery, and retain a route back to the ordinary workflow.
Decisions outside the loop
Make the operational exceptions visible.
These rules connect the example to everyday support work. They need explicit product behavior, not just instructions in a system prompt.
- Approval is attached to an effect
- Show the destination, full note, source context, and unresolved issues together. Rejecting a proposal must end that path. An edited proposal returns to review rather than inheriting an earlier approval.
- Concurrency is part of correctness
- A read followed by a write can race with another worker. Where supported, use a server-enforced conditional write such as If-Match, or an equivalent transaction/version check. A client-side comparison alone does not close that gap.
- Cancellation has a boundary
- Stopping a run can prevent future calls but may not undo a dispatched write. Show any in-flight operation and reconcile it. A compensating action has its own permissions and consequences; it is not an automatic rollback.
- Evidence has its own access rules
- Operators need event identifiers, policy decisions, proposal versions, and receipts. They do not need every sensitive source copied into every log. Define access, retention, redaction, and incident evidence needs for each record.
Implementation questions
Resolve these before choosing the runtime.
The useful comparison is whether your team can implement, inspect, test, and operate the required behavior.
- Do we need an agent framework?
- Not necessarily. A small explicit loop can be easier to inspect for one bounded task. A framework may help with persistence, tool orchestration, or tracing, but verify its retry, cancellation, authorization, and versioning behavior against your own contracts.
- Should the first version use several agents?
- Only if a measured task need justifies the extra coordination. Start with one accountable run and a limited tool set. Additional agents introduce more handoffs, context boundaries, and failure states; agreement between models does not establish authority.
- Does a tool schema make a call safe?
- No. It can establish shape, not permission, factual correctness, or the acceptability of an effect. A correctly formed call can still target the wrong record. Validate semantics and authorization in trusted code at the execution boundary.
- What if the system has no idempotent write API?
- Do not promise duplicate-free retries by adding a local key alone. Determine whether the downstream system can recognize an operation or expose conclusive status. Where ambiguity cannot be resolved, retain an operator-controlled path and keep the outcome explicitly unknown.
- What should a release report contain?
- The tested task scope, dataset and version details, success and failure definitions, observed failures, reviewer workload, resource use, recovery results, owners, and unresolved limits. A single aggregate score is too little to decide which actions may run.
Source basis
Sources behind the control model.
- 01
Anthropic
Building effective agentsVendor engineering perspective on workflows, agents, simple architectures, and tool interfaces. The article notes that its tooling landscape has changed since 2024.
- 02
Anthropic
Demystifying evals for AI agentsSupports evaluating trajectories and outcomes across repeated tasks rather than judging final prose alone.
- 03
OWASP
AI Agent Security Cheat SheetSecurity guidance for tool scope, untrusted context, memory, approvals, and abuse-case testing. Guidance is not evidence that a specific implementation is secure.
- 04
OWASP
Authorization Cheat SheetBasis for enforcing least privilege and checking access at the application boundary, independently of model output.
- 05
Amazon Builders' Library
Making retries safe with idempotent APIsExplains ambiguous outcomes, request identity, duplicate effects, and the importance of the service's idempotency contract. The ticket trace is Werkon's illustrative application.
- 06
IETF
RFC 9110: HTTP SemanticsDefines If-Match and conditional request semantics. An integration must actually support and enforce the relevant precondition.
- 07
UK National Cyber Security Centre
Guidelines for secure AI system developmentLifecycle guidance spanning design, development, deployment, and operation. It does not certify the architecture described here.
Start with one real workflow
A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.
Show Us the WorkflowStart with the free automation readiness checklistOBSERVEQUANTIFYDECIDEBUILD
