Decision guide / bounded AI trials
A low-risk AI project earns the label through its boundary, not its topic.
No AI project is low risk merely because it drafts, searches, classifies, or stays internal. A sensible first trial limits who and what it can affect, keeps consequential authority with people, measures against the current process, and can be stopped without leaving a hidden dependency behind.
The short answer
Start where a wrong output cannot quietly become a real action.
For this guide, a comparatively lower-risk first AI project has a narrow purpose, known users, approved data, no hidden authority, a measurable baseline, observable failure, a working non-AI fallback, and an exit that has been tested. Begin with no live influence when possible. Add exposure only after the prior boundary has produced useful local evidence.
- Task names do not determine risk
- Summarization can expose confidential material or distort a professional conclusion. Classification can redirect access or opportunity. Drafting can create an external promise. Assess the actual context and consequence, not the friendly label.
- Human review is a system property
- Review is credible only when a named person has enough evidence, competence, time, authority, and a usable override. A nominal approval step can increase risk when volume, interface, or incentives make scrutiny unrealistic.
- Reversible means operationally reversible
- A stop button is not enough. The organization must preserve the prior workflow, remove access, prevent further writes, reconcile changed records, retain required evidence, notify affected owners, and continue service without the model or provider.
Six controlled trial shapes
Choose the least exposed lane that can still answer the decision.
These are trial shapes, not a maturity ladder and not claims of safety. A team can learn in one lane and decide never to move further. Each shape should be rejected when its stated boundary cannot be enforced or when the decision can be answered more simply without AI.
Protected benchmark only
01Decision under test: Can any candidate method perform the bounded task well enough on protected cases to justify a more realistic trial?
- Entry condition
- The target task, expected outputs, error categories, and evaluation cases can be defined before the candidate sees them, and no live workflow access is required.
- Required evidence
- A written task contract, protected representative cases, expected evidence, segment labels, simple non-AI baseline, scoring rules, prohibited outputs, evaluator instructions, and access and retention controls.
- Exposure boundary
- Run isolated candidates against fixed test material. Do not expose outputs to operational users, connect production tools, update records, contact anyone, or tune on the protected cases.
- Retained authority
- Domain owners define the expected result and material errors. Evaluation owners protect the holdout and decide whether the evidence is interpretable. No model output controls a business decision.
- Measures
- Results by case type and consequence, comparison with the baseline, abstention and refusal behavior, unsupported output rate, evaluator disagreement, run cost, latency, repeatability, and failed-case review.
- Escalation risk
- A convenient test set can exclude rare cases, encode one reviewer, leak into prompts or tuning, reward superficial similarity, or measure a proxy that has little connection to the real workflow.
- Reversal plan
- Revoke candidate access, retain the evaluation record, delete temporary material under the approved rule, and continue the current process unchanged. A failed benchmark creates no live migration to undo.
Historical replay lane
02Decision under test: Would the candidate have supplied useful evidence on past work without changing the decisions that were actually made?
- Entry condition
- Past inputs, timestamps, decisions, corrections, and outcomes can be reconstructed lawfully, and the team can prevent future information from leaking into the replay.
- Required evidence
- Time-correct historical records, original decisions and reasons, later outcomes kept separate, versioned policies, known interventions, missing-data markers, permissions, and a replay protocol.
- Exposure boundary
- Recreate the information available at each historical point, run the candidate without writeback, compare it with the recorded process, and keep retrospective outcome data outside the candidate context.
- Retained authority
- Process owners determine whether the reconstruction is faithful, whether the original decision remains a valid comparator, and which discrepancies deserve investigation rather than automatic correction.
- Measures
- Coverage, error and abstention by time and segment, difference from the original process, corrected versus disputed records, reviewer effort, data gaps, stability across policy versions, and plausible effect on the target decision.
- Escalation risk
- Hindsight leakage, changed policy, incomplete records, survivor bias, model exposure to later outcomes, a historically weak comparator, or confident conclusions from a period unlike current work.
- Reversal plan
- Remove replay data and access as required, preserve the reproducible comparison, and leave current operations untouched. Any next step requires a new boundary and approval rather than inheriting replay permission.
Live shadow comparison
03Decision under test: Can the candidate handle current variation and timing while the existing workflow remains the only path that affects people or records?
- Entry condition
- Inputs can be copied safely, shadow processing cannot delay or alter live work, outcomes can be joined later, and the organization can afford the observation period and review burden.
- Required evidence
- Current permitted inputs, exact version identity, routing and timing data, baseline decisions, delayed outcome joins, failure injection cases, service limits, privacy rules, and an independent shutdown path.
- Exposure boundary
- Run beside the live process. Hide candidate outputs from decision-makers when comparison bias matters, block all write and communication paths, and record candidate results for later matched analysis.
- Retained authority
- Operational owners keep full control of the real decision and service. Technical owners can isolate or stop the shadow lane without touching production, and evaluators decide when comparison evidence is sufficient.
- Measures
- Match rate, disagreement, error by consequence and segment, abstention, latency, availability, input drift, operating cost, failure containment, and whether shadow results can be reconciled with final outcomes.
- Escalation risk
- A shadow connection can still expose data, consume scarce capacity, create a covert dependency, influence users through leaked output, or produce misleading results when the live decision changes the observed outcome.
- Reversal plan
- Disable the copied feed and compute, verify that no write path exists, remove shadow credentials, reconcile retained records, and confirm the original workflow was never dependent on the candidate.
Read-only evidence assistant
04Decision under test: Can a small group find and inspect permitted evidence more effectively without letting generated text become an authoritative answer?
- Entry condition
- The corpus is bounded and maintained, document and user permissions are enforceable, sources can be shown at passage level, and qualified users can judge support and completeness.
- Required evidence
- Approved source inventory, document versions, access rules, source owners, representative questions, expected citations, known gaps, forbidden topics, review rubric, and a reliable manual search path.
- Exposure boundary
- Retrieve candidate passages and compose a clearly non-authoritative response with citations. Permit no source changes, external communication, record updates, tool use, or action based solely on the response.
- Retained authority
- Users inspect the cited source and own any later interpretation or action. Source owners correct or withdraw material, and security and privacy owners control access and incident response.
- Measures
- Citation correctness and coverage, unsupported claims, missed evidence, permission isolation, useful abstention, user corrections, search and review time, over-reliance signals, and unresolved-question volume.
- Escalation risk
- Stale authority, retrieval misses, conflicting documents, hidden instructions in source material, permission leakage, a fluent answer replacing source inspection, or later reuse outside the tested context.
- Reversal plan
- Disable the interface and retrieval credentials, remove indexes and caches under the retention rule, preserve issue records, and return users to the maintained source library and manual search process.
Draft-only workbench
05Decision under test: Can a candidate prepare a useful internal draft while an accountable author retains every claim, recipient, commitment, and send decision?
- Entry condition
- The material is low consequence, required facts and tone rules are available, the final author has real review capacity, and technical controls can prevent autonomous sending or publication.
- Required evidence
- Permitted facts and sources, approved audience and purpose, required disclosures, prohibited claims, privacy rules, examples with usage rights, author rubric, review time, and a manual drafting baseline.
- Exposure boundary
- Create a visibly marked draft in an isolated workspace. Run deterministic checks, show sources and unresolved items, and require the author to move approved content through a separate communication path.
- Retained authority
- The author chooses the recipient, verifies facts, owns professional judgment and tone, rejects or rewrites content, and initiates any send or publish action outside the model-controlled environment.
- Measures
- Unsupported and prohibited claims, omissions, privacy findings, corrections, rejection rate, review time, author disagreement, copied phrasing, near-miss incidents, and whether the draft improved the whole task rather than typing speed alone.
- Escalation risk
- Review fatigue, fabricated support, sensitive context in prompts, inappropriate reuse, concealed automation, wrong audience, accidental send paths, or fluency that turns approval into a ritual.
- Reversal plan
- Remove generation access and integrations, retain approved records only as policy requires, restore the manual template, and verify that no scheduled job, token, webhook, or user habit depends on the workbench.
Sampled assistive queue
06Decision under test: Can a bounded suggestion improve a reversible review step for a small population without hiding missed cases or granting entitlement, priority, or final disposition?
- Entry condition
- The queue has stable owners, a correctable intermediate decision, measurable outcomes, deliberate control or holdout cases, capacity to review both suggestions and misses, and a safe manual route.
- Required evidence
- Current taxonomy, approved population, sampling rule, baseline queue data, rare and protected cases, human review standard, override reasons, outcome definition, capacity limits, and an immediate disable path.
- Exposure boundary
- Offer one suggestion inside a capped queue, apply policy and permission rules deterministically, require reviewer disposition, preserve a control sample, inspect non-suggested cases, and prevent downstream action without approval.
- Retained authority
- Queue owners define categories and service rules. Reviewers decide the case. Accountable leaders set exposure, approve changes, respond to unfair or harmful patterns, and can return all work to the manual path.
- Measures
- Outcome and handling time against the baseline, errors by class and consequence, missed cases from sampling, override patterns, queue balance, reviewer load, affected-group differences, incidents, and downstream reconciliation.
- Escalation risk
- The suggestion can anchor reviewers, shift service quality, hide false negatives, amplify historical practice, overload one queue, expand beyond the approved population, or become a de facto decision despite formal review.
- Reversal plan
- Set exposure to zero, route all items through the manual rule, remove model access, reconcile pending and completed cases, notify queue owners, and review whether any policy or staffing decision had begun to rely on the suggestion.
Exposure matrix
A first trial should state its ceiling before anyone sees a result.
The matrix compares operating exposure, not universal safety. Any row can be unsuitable when it uses sensitive data, affects protected or vulnerable people, enters a regulated or safety-critical context, or lacks a credible owner, baseline, fallback, and exit.
| Trial shape | Exposure ceiling | Decision retained | Evidence for next gate | Exit trigger |
|---|---|---|---|---|
| Protected benchmark | Fixed approved cases; no operational user, connection, write, or action. | Candidate produces test output only; evaluators own interpretation. | Performance and failure by case type versus a simple baseline. | Task or expected result cannot be defined, protected, or scored meaningfully. |
| Historical replay | Past time-correct inputs; no current workflow influence or writeback. | Process owners validate reconstruction and comparison meaning. | Segmented errors, disagreement, reviewer effort, and data gaps over time. | Future leakage, missing provenance, or policy change makes the comparison invalid. |
| Live shadow | Copied current inputs; output hidden; production remains authoritative. | Operational owners keep the real path; evaluators inspect matched results. | Current variation, timing, containment, cost, and delayed outcome evidence. | Shadow work can affect service, expose data, or cannot be reconciled reliably. |
| Read-only evidence | Small approved user group; bounded corpus; no source or system changes. | User verifies source support and owns every later interpretation or action. | Citation support, permission isolation, abstention, corrections, and review time. | Source authority or access cannot be enforced, or answers replace source inspection. |
| Draft-only workbench | Marked internal draft; no autonomous recipient, send, publish, or commitment. | Named author owns facts, judgment, audience, wording, and release. | Unsupported claims, corrections, rejection, privacy findings, and total task effort. | Review becomes ceremonial or a technical path can release unapproved content. |
| Sampled assistive queue | Capped population; one reversible suggestion; holdout and manual path preserved. | Reviewer decides each case; owners retain policy, priority, and service authority. | Outcomes, misses, overrides, workload, group differences, and incidents. | Anchoring, unfair service, hidden misses, overload, or scope expansion exceeds the gate. |
Five decision gates
Reduce exposure before you ask whether the model is impressive.
The sequence starts with consequences and ends with a recorded decision. It does not assume that a successful test should become a live feature, or that more exposure is the only form of progress.
- 01
Screen the consequence
Name affected people, rights, safety, money, access, work, communication, records, data, jurisdiction, and professional duties. Exclude prohibited, high-consequence, or unowned uses from a casual first trial.
- 02
Fix the control envelope
Write allowed users, inputs, outputs, tools, actions, systems, population, volume, duration, permissions, human decisions, logging, fallback, stop authority, and deletion or retention before connecting anything.
- 03
Record the baseline
Measure the current process, including quality, time, exceptions, review effort, delay, cost, incidents, and outcomes. State the hypothesis and thresholds without converting uncertain benefits into a promise.
- 04
Test the least exposed lane
Prefer benchmark, replay, or shadow evidence before visible use. Exercise refusal, wrong data, changed policy, unavailable provider, access failure, overload, correction, fallback, shutdown, and recovery.
- 05
Close the gate
Compare the baseline, failure distribution, human workload, incidents, operating cost, control evidence, and exit test. Record expand, revise, hold, or stop with a named owner and a new boundary for any next step.
Control tests
Four questions decide whether the boundary is real.
A diagram or policy statement does not contain a system. Before the trial begins, demonstrate each control with a failure case and identify the person who owns the response.
- Who and what can be affected?
- List users, non-users, subjects, groups, records, services, and downstream consumers. Separate direct effects from plausible indirect effects, and prevent unapproved reuse, population growth, or secondary decisions.
- What can the model cause?
- Enumerate read, suggest, draft, rank, route, write, send, approve, deny, purchase, schedule, and delete paths. Enforce identity, permission, policy, limits, and consequential authority outside model discretion.
- Can failure be seen and corrected?
- Make source, version, uncertainty, override, abstention, incident, reviewer load, missed-case sampling, and downstream outcome observable within approved privacy limits. Provide correction and appeal where people can be affected.
- Can the organization leave cleanly?
- Keep the prior path viable. Test access revocation, shutdown, rollback, record reconciliation, data export and deletion, provider removal, retained evidence, owner notification, and service continuity before dependency grows.
Questions before approval
Short answers that keep low risk from becoming a slogan.
The answer changes with context, so each response points back to evidence and authority the organization must supply. A project should stay private when those facts are missing.
- Are internal chatbots automatically low risk?
- No. An internal tool can expose restricted data, present stale policy as authority, influence employment or financial work, send information to a provider, or create over-reliance. Judge the corpus, users, action, affected people, access, review, consequence, and exit.
- What makes an AI trial reversible?
- The prior process still works, the candidate can be isolated without service loss, access can be revoked, pending work can be reconciled, changed records can be identified and corrected, required evidence remains available, and provider-held data can be handled under an approved exit rule.
- How long should a first trial run?
- There is no responsible universal duration. Run long enough to observe the relevant volume, segments, rare and costly cases, operational cycles, reviewer load, drift, failures, and delayed outcomes. Stop earlier when a hard boundary fails or the evidence cannot answer the decision.
- Which metric proves the project is worth continuing?
- No single metric can do that. Compare task quality, error distribution, abstention, human effort, delay, cost, incidents, affected-group results, operational reliability, and realized outcomes with the current process. Weight them by local consequence and uncertainty.
- When may a shadow trial become visible to users?
- Only after the shadow evidence meets its written gate, critical failures and controls have been tested, users and affected people are considered, review capacity is credible, fallback and exit work, and accountable owners approve a new limited exposure. Shadow success is not production approval.
Source basis
Sources behind the control model.
- 01
National Institute of Standards and Technology
Artificial Intelligence Risk Management Framework 1.0The 2023 voluntary, non-sector-specific framework organizes AI risk work around Govern, Map, Measure, and Manage. It supports contextual risk framing and explicit removal options, but it does not label a local project low risk or establish a result.
- 02
National Institute of Standards and Technology
AI Risk Management Framework program pageThe live program page says AI RMF 1.0 is being revised in 2026 and remains voluntary. It is included so the guide exposes the current revision state instead of treating the 2023 publication as fixed guidance.
- 03
National Institute of Standards and Technology
AI Resource CenterThe live resource center supports operational use of the AI RMF and access to testing, evaluation, verification, and validation material. Its listed tools and resources are aids, not NIST endorsement or proof that a chosen measurement fits a local use.
- 04
National Institute of Standards and Technology
TEVV-Athlon Framework for Evaluating AI SystemsThe August 2026 page presents NIST AI 200-2 as an initial public draft and requests comment through October 6, 2026. Its customizable assessment framing is current research input, not final guidance or validation of this guide's trial shapes.
- 05
National Institute of Standards and Technology
Assessing Risks and Impacts of AI pilot evaluation reportThe 2025 report documents one pilot involving seven submitted applications, three scenarios, and model, red-team, and field testing. It illustrates layered evaluation and human testing, but its small pilot does not establish universal measures or outcomes.
- 06
National Institute of Standards and Technology
Generative AI Profile for the AI RMFThe 2024 cross-sector profile describes risks that generative AI can create or intensify and suggests actions across governance, mapping, measurement, and management. It is voluntary and cannot replace a local impact, legal, or operational assessment.
- 07
National Institute of Standards and Technology
AI RMF PlaybookThe playbook offers voluntary suggested actions for the AI RMF functions. Its page was updated in June 2026 and says it will change after the framework revision, so selections remain provisional and context-specific.
- 08
National Institute of Standards and Technology
SP 800-218A secure development profile for generative AIThe final 2024 community profile adds generative AI and dual-use foundation model practices for producers and acquirers. It is intended to work with the base Secure Software Development Framework and is not complete security assurance.
- 09
National Institute of Standards and Technology
Privacy FrameworkThe framework supports enterprise privacy risk management and does not have the force of law. It helps expose data processing and affected-person risk, but it cannot determine local lawfulness, rights, or acceptable use.
- 10
National Institute of Standards and Technology
Cybersecurity Framework 2.0The 2024 framework supplies a non-prescriptive taxonomy for governing and communicating cybersecurity outcomes. It can structure access, protection, detection, response, and recovery questions without proving that an AI trial is secure.
- 11
U.S. Government Accountability Office
Artificial Intelligence Accountability FrameworkThe 2021 federal framework groups accountability practices under governance, data, performance, and monitoring and provides questions for managers and assessors. Its government context informs inquiry but does not determine commercial risk or assurance.
- 12
UK Government
Guidelines for AI procurementThe 2020 public-sector guide is explicitly not exhaustive. Its multidisciplinary planning, data assessment, oversight, ongoing testing, knowledge transfer, lifecycle cost, and end-of-life questions are useful procurement context, not current universal law or a private-sector approval method.
- 13
UK Department for Science, Innovation and Technology
Introduction to AI assuranceThe 2024 introduction describes assurance as measuring, evaluating, and communicating trustworthiness with techniques proportionate to context. It is introductory guidance and says multiple techniques are needed; it does not certify any project or define low risk universally.
- 14
UK Department for Science, Innovation and Technology
Portfolio of AI assurance techniquesThe living portfolio maps examples and assurance techniques across lifecycle stages. The page explicitly says included case studies are not government endorsements, so it is evidence of available technique types rather than proof of provider or project quality.
- 15
European Commission
AI Act regulatory frameworkThe current page describes prohibited, high, transparency, and minimal or no risk categories and a staged application timeline updated through 2026. Classification depends on the exact use, actor, date, and law, so this guide does not convert the categories into legal advice.
- 16
Information Commissioner's Office
AI and data protection risk toolkitThe UK regulator's toolkit addresses risks to individual rights and freedoms. Its page says the guidance is under review because of the Data (Use and Access) Act, so it must not be treated as settled or universal compliance guidance.
- 17
Organisation for Economic Co-operation and Development
Framework for the Classification of AI SystemsThe framework links technical and contextual characteristics with potential effects on individuals, society, and the planet. It supports more precise classification discussion but remains a generic policy tool, not a local risk rating or approval.
- 18
Organisation for Economic Co-operation and Development
Explanatory memorandum on the updated definition of an AI systemThe 2024 memorandum explains the OECD definition adopted for its AI Recommendation. This guide uses its inference-centered boundary to distinguish AI from deterministic software; other laws and standards can use different definitions.
- 19
International Organization for Standardization
ISO/IEC 42001:2023 AI management systemsThe public record describes requirements for an organizational AI management system and continuing improvement. The complete standard is paid material, and the record does not establish certification, conformity, or effective operation for an organization.
- 20
International Organization for Standardization
ISO/IEC 23894:2023 AI risk management guidanceThe public record describes customizable guidance for integrating AI-specific risk management into organizational activity. The complete standard is paid and does not classify a local trial, set risk acceptance, or prove compliance.
- 21
UK National Cyber Security Centre
Guidelines for secure AI system developmentThe multi-agency guidance addresses secure design, development, deployment, operation, and maintenance for AI system providers, including those using hosted models and external APIs. Local threats, architecture, duties, and control evidence still require assessment.
- 22
Microsoft HAX Toolkit
Guidelines for Human-AI InteractionThe provider toolkit presents research-based interaction guidance for initial use, ongoing interaction, error, and change over time. It is design input, not a universal interface recipe or proof of local usability, accessibility, or safe human review.
- 23
World Wide Web Consortium
Web Content Accessibility Guidelines 2.2The W3C Recommendation supplies testable web-content accessibility criteria for human-facing interfaces. Conformance applies to complete page variations and needs appropriate evaluation; citing it does not establish product accessibility.
Start with one real workflow
A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.
Show Us the WorkflowStart with the free automation readiness checklistOBSERVEQUANTIFYDECIDEBUILD
