Decision guide / AI projects
Six AI projects worth testing before you scale anything.
A useful AI project is not a model looking for somewhere to land. It is a bounded change to a real workflow, with named inputs, human authority, failure handling, and evidence that can support a decision to continue, change course, or stop.
The short answer
Start with a decision and a controlled way to learn.
For this guide, an AI project is a bounded effort to introduce, change, or evaluate a system that infers from inputs how to produce predictions, content, recommendations, or decisions. It should be tied to one real workflow and one named business decision. If ordinary rules can produce the required result reliably, use ordinary automation instead.
- AI is not every automation
- A deterministic rule, calculation, validation, or workflow can be the safer and more inspectable choice. Use model inference only where variability or ambiguity makes it useful, and keep rules around it where the business already knows the answer.
- A model is not the system
- The operating system also includes source data, access controls, interfaces, deterministic checks, human review, logs, monitoring, incident handling, change control, recovery, and an owner who can withdraw it.
- A demonstration is not evidence
- A polished example can show possibility. It cannot establish performance across real cases, safe failure, manageable review load, production reliability, legal fit, or a business result. Those need local evaluation over time.
Six project patterns
Compare the work around the model, not just the output.
These patterns can use different models and products. Their useful difference is the job they perform, the evidence available to check them, and the authority they are allowed to exercise. None is automatically a good first project.
Knowledge and evidence assistant
01Decision: Which source-supported information should a reviewer use to answer a question, prepare a case, or continue an investigation?
- Best when
- The organization has a bounded, permissioned body of current material and reviewers can judge whether an answer is supported, complete, and appropriate for the situation.
- Inputs
- Approved source documents, source owners, access rules, version and freshness data, known gaps, representative questions, expected evidence, and explicit subjects the system must not answer.
- Workflow
- Retrieve candidate passages, show their provenance, compose a bounded answer, expose uncertainty or missing evidence, and route the result to a person before any consequential use or writeback.
- Human authority
- People approve external or consequential answers, resolve conflicting sources, change access, accept exceptions, and decide whether the underlying record is authoritative.
- First evidence
- A protected set of real questions with expected evidence, citation correctness and coverage, unsupported-claim review, abstention behavior, permission tests, reviewer corrections, and review time.
- Failure modes
- Stale or contradictory documents, retrieval misses, access leakage, hidden instructions in source material, fabricated support, omitted context, over-reliance, and a feedback queue nobody owns.
- Operating change
- The source library becomes a maintained product. Someone must own freshness, permissions, disputed answers, feedback triage, incident response, and the decision to remove a source or suspend the assistant.
Document intake and validation
02Decision: How should an incoming document be represented, which fields are sufficiently supported, and which exceptions need a person?
- Best when
- Documents are frequent enough to justify a pipeline, the target record and business rules are explicit, and field-level evidence can remain attached to every accepted or rejected value.
- Inputs
- Representative files and images, document and field taxonomies, target schemas, field definitions, deterministic rules, duplicate policy, labeled exceptions, privacy constraints, and downstream acceptance criteria.
- Workflow
- Identify the document, extract candidate fields with page or region provenance, validate type and business rules deterministically, detect duplicates, and send ambiguity or contradiction to an exception queue.
- Human authority
- People resolve ambiguous evidence, approve material records, correct classifications, authorize downstream actions, and decide how to treat handwriting, missing pages, conflicting values, and unsupported fields.
- First evidence
- Field-level accuracy by document and field type, false acceptance, false rejection, provenance coverage, duplicate behavior, exception volume, correction patterns, latency, and downstream reconciliation.
- Failure modes
- Image quality, layout drift, handwriting, copied identifiers, unit mistakes, prompt injection inside documents, personal-data exposure, silent default values, and apparent confidence without usable evidence.
- Operating change
- Intake becomes a monitored queue rather than a one-time upload. Schema versions, source retention, reviewer workload, correction capture, reprocessing, and downstream reconciliation all need explicit owners.
Request classification and queue routing
03Decision: Which queue, category, priority, or next review should receive a request without allowing a model to grant access or make the underlying business decision?
- Best when
- Categories have operational meaning, owners and service paths are stable enough to maintain, and a wrong route can be detected and corrected before material harm occurs.
- Inputs
- A current taxonomy, queue ownership, labeled historical requests, rare and high-consequence cases, priority rules, prohibited actions, identity and permission context, and an escalation path.
- Workflow
- Normalize the request, suggest labels and a route, apply deterministic permission and priority policy, expose uncertainty, send protected classes to review, and record the final route and any override.
- Human authority
- People define categories and priority policy, approve sensitive routing, handle conflicts, change ownership, and retain authority over entitlement, eligibility, discipline, credit, care, and other consequential outcomes.
- First evidence
- Confusion by class, rare-case recall, high-consequence misroutes, time to correction, override reasons, queue balance, policy-block behavior, drift in category volume, and end-to-end resolution rather than label accuracy alone.
- Failure modes
- Historical labels encode inconsistent practice, category drift, urgency hidden in free text, strategic manipulation, biased routing, permissions inferred from content, automation loops, and unstaffed queues.
- Operating change
- The taxonomy becomes an operational contract. Queue owners need a way to correct routes, update definitions, review drift, identify new classes, and stop automation when ownership or policy changes.
Drafting and response preparation
04Decision: Which draft, evidence, and unresolved questions should an accountable person review before communicating or committing anything?
- Best when
- The final author has real review time and subject authority, source and tone rules can be stated, and the draft can be withheld when evidence or context is missing.
- Inputs
- Approved facts and sources, recipient context, purpose, tone and policy rules, required disclosures, prohibited claims, examples with usage rights, privacy constraints, and a review rubric.
- Workflow
- Assemble permitted context, produce a draft with visible source support and unresolved items, run deterministic policy and data checks, then require review before sending, publishing, or changing a system of record.
- Human authority
- A person owns factual accuracy, professional judgment, promises, tone, recipients, disclosure, and the send or publish action. Approval is a real decision, not a button pressed after a cursory glance.
- First evidence
- Factual support, prohibited-claim detection, reviewer corrections, rejected drafts, unresolved-item capture, inappropriate disclosure, time spent reviewing, and incidents. Edit distance can describe workload, not quality by itself.
- Failure modes
- Plausible unsupported statements, private data in context or output, copied phrasing, policy drift, fabricated references, wrong recipients, hidden reviewer fatigue, and fluency that suppresses skepticism.
- Operating change
- Source, policy, and disclosure rules need versioning. Review responsibility, correction feedback, incident handling, retained records, and the ability to disable a template or channel become part of the communication process.
Forecasting and exception review
05Decision: Which future range or unusual condition deserves investigation, and what should an accountable planner decide after seeing it?
- Best when
- The outcome and forecast horizon are measurable, a simple baseline exists, error costs can be discussed, and planners can act on exceptions without treating a forecast as an order or commitment.
- Inputs
- Time-stamped outcomes, known drivers, horizon and refresh cadence, segment definitions, lead times, decision constraints, missing-data rules, naive baselines, error costs, and recorded interventions.
- Workflow
- Create a time-correct comparison, produce ranges and exception signals, compare them with the baseline, expose missing or changed conditions, send exceptions to a planner, and return actual outcomes for review.
- Human authority
- People approve purchases, staffing, pricing, capacity, financial treatment, safety actions, and other commitments. They can reject a signal, record outside knowledge, and halt use after a regime change.
- First evidence
- Backtesting by horizon and segment, error against a simple baseline, interval calibration, false alarms, missed costly events, stability after interventions, data delay, decision use, and realized outcomes after enough time passes.
- Failure modes
- Future information leaking into training, broken timestamps, regime shifts, unrecorded interventions, sparse segments, correlated errors, misleading averages, false precision, and feedback from prior model-influenced decisions.
- Operating change
- Planning moves toward exception review with recorded reasons. Data delays, overrides, external events, model changes, error cost, and actual results need a recurring review rather than a one-time accuracy report.
Visual inspection support
06Decision: Which image, frame, region, or item should a trained person inspect, and what evidence should accompany that review?
- Best when
- Capture conditions can be controlled or measured, the target condition is defined, representative images can be lawfully retained, and a false negative has a safe review or sampling path.
- Inputs
- A defect or condition taxonomy, representative images across cameras and environments, region labels, capture requirements, acceptance tolerances, rare and borderline cases, privacy rules, and the downstream disposition process.
- Workflow
- Check capture quality, preserve the source image, produce a candidate region and score, apply a bounded threshold, route uncertain and sampled normal cases to inspection, and record the final human disposition.
- Human authority
- Qualified people define the inspected condition, approve disposition, stop production or service when required, set sampling policy, resolve disputed labels, and decide whether capture or model changes are acceptable.
- First evidence
- Performance by condition, product, site, camera, lighting, angle, subgroup, and time; false negatives; false positives; image-quality rejection; localization quality; reviewer agreement; latency; drift; and sampled misses.
- Failure modes
- Lighting and camera shifts, shortcut learning, mislabeled or unrepresentative images, rare unseen conditions, privacy exposure, unsuitable compression, latency, reviewer over-trust, and a score used beyond its evaluated context.
- Operating change
- Cameras, capture instructions, sampling, label review, privacy, model versions, false-negative investigation, maintenance, and safe fallback become part of the inspection process, not background technical details.
Decision matrix
Know what a first test must prove and what it must never decide.
The table is a starting point for discovery, not a risk ranking. The same pattern can move from modest to unacceptable risk when its data, affected people, action, scale, or failure consequence changes.
| Project pattern | First useful evidence | AI boundary | Human authority | Stop signal |
|---|---|---|---|---|
| Knowledge assistant | Supported answers and useful abstention on protected real questions. | Retrieve and compose candidates from permitted sources. | Approve consequential answers and source authority. | Sources cannot be governed or access boundaries cannot be enforced. |
| Document intake | Field-level results with provenance and honest exception volume. | Classify and extract candidates before deterministic validation. | Resolve ambiguity and approve material records or actions. | False acceptance or reviewer load exceeds the safe process. |
| Queue routing | Rare-class and high-consequence route performance with overrides. | Suggest category and route inside policy constraints. | Own categories, protected cases, priority, and entitlement. | No stable owner exists for categories, corrections, or destination queues. |
| Draft preparation | Source support, corrections, rejected drafts, and real review effort. | Prepare a reviewable draft and identify missing context. | Own claims, judgment, recipient, and send or publish action. | Review is ceremonial or unsupported claims are hard to detect. |
| Forecast review | Time-correct backtest against a simple baseline by useful segment. | Estimate ranges and highlight exceptions. | Approve commitments and respond to changed conditions. | The outcome, horizon, baseline, or cost of error cannot be defined. |
| Visual inspection | Condition-specific errors across representative capture contexts. | Suggest regions, conditions, or review priority. | Approve disposition, sampling, and production or safety action. | Representative images or a safe path for missed conditions are unavailable. |
Selection method
Turn the idea into a decision record before funding a build.
A first project should make uncertainty cheaper to resolve. These steps keep the model choice downstream of the work, evidence, and authority that determine whether the project is useful.
- 01
Name the decision
Write the user, trigger, decision, current method, accountable owner, affected people, prohibited actions, and consequence of a wrong or late result.
- 02
Measure today
Establish volume, time, error, exception, review load, delay, and outcome evidence. Define a simple non-AI baseline and the costs that matter locally.
- 03
Prove the inputs
Confirm provenance, rights, permissions, quality, representativeness, retention, freshness, labels, and the owners who can correct or withdraw data.
- 04
Run a bounded test
Start in replay, shadow, or restricted assistive use. Protect evaluation cases, capture failure and human effort, and keep rollback independent of the model.
- 05
Choose by evidence
Compare the baseline, error distribution, review burden, operating cost, security and privacy findings, incidents, and ownership readiness. Continue, change, or stop explicitly.
Operating controls
Scaling means scaling ownership as well as use.
The right control depends on context. At minimum, a production proposal should make these four responsibilities inspectable before more people, data, or actions enter the system.
- Authority and permissions
- Keep identity, access, allowed tools, data boundaries, prohibited actions, approvals, and separation of duties outside model discretion. Name who can authorize, override, suspend, and withdraw each use.
- Evaluation and release
- Bind each released version to protected cases, task and segment measures, security and privacy review, accessibility checks where people interact, known limits, acceptance authority, and a change trigger.
- Observation and recovery
- Record inputs and outputs within approved privacy limits, decisions, overrides, failures, latency, cost, version, and downstream reconciliation. Provide incident handling, safe fallback, rollback, and correction paths.
- Data, supplier, and exit
- Track data provenance and retention, component and provider terms, location and sub-processors where relevant, dependency changes, portability, deletion, replacement, and how the organization continues if a model or provider is removed.
Questions leaders usually ask next
Short answers that preserve the hard part.
These answers are deliberately conditional. A project becomes specific only when the organization supplies its workflow, evidence, affected people, authority, and failure consequences.
- Which pattern is lowest risk?
- A reversible, assistive use with bounded data, visible evidence, no external action, skilled review, and a safe non-AI fallback can be easier to test. It is not automatically low risk. Assess the actual use, affected people, data, scale, jurisdiction, and consequence of error.
- How should we estimate return on investment?
- Measure the current process first, then include implementation, integration, data work, evaluation, review, correction, incidents, platform use, maintenance, and exit. Compare observed value and error cost after a bounded trial. Do not turn a model demo or generic benchmark into a financial promise.
- Should we buy a product or build a system?
- Decide from workflow fit, data and access boundaries, integration, evaluation access, controllability, observability, contractual terms, supplier change, portability, and exit. A product can reduce construction work; it does not remove local governance or operating responsibility.
- What separates a prototype from production?
- Production needs real ownership, representative evaluation, access control, change and release records, usable human review, monitoring, support, incident response, correction, recovery, cost control, supplier management, and retirement. More traffic does not supply those capabilities.
- Is a human in the loop enough?
- Only if the review is defined. Name the reviewer, evidence shown, time available, competence required, decision retained, override path, workload limit, record created, and response when review fails. Otherwise the phrase hides rather than manages responsibility.
Source basis
Sources behind the control model.
- 01
National Institute of Standards and Technology
Artificial Intelligence Risk Management Framework 1.0The 2023 framework organizes voluntary, non-sector-specific AI risk work for organizations that design, develop, deploy, or use AI. Its risk and trustworthiness structure informs the guide, but it does not prescribe one project or establish a local result.
- 02
National Institute of Standards and Technology
AI Risk Management Framework program pageThe live program page says AI RMF 1.0 is being revised in 2026 and continues to describe it as voluntary. It is included so readers can see the current revision state rather than treating the 2023 document as fixed guidance.
- 03
National Institute of Standards and Technology
Generative AI Profile for the AI RMFThe 2024 cross-sector profile describes risks that generative AI can create or intensify and suggests actions across governance, mapping, measurement, and management. It is voluntary guidance, not a complete local risk assessment or legal rule.
- 04
National Institute of Standards and Technology
AI RMF PlaybookThe playbook offers voluntary suggestions for the Govern, Map, Measure, and Manage functions. Its page was updated in June 2026 and says the playbook will change after the AI RMF revision, so selections need current local review.
- 05
National Institute of Standards and Technology
SP 800-218A secure development profile for generative AIThe final 2024 community profile adds generative AI and dual-use foundation model practices for model producers, system producers, and acquirers. It is intended to be used with the base Secure Software Development Framework, not as standalone assurance.
- 06
National Institute of Standards and Technology
Privacy FrameworkThe framework supports enterprise privacy risk management and explicitly does not have the force of law. It helps frame data processing and affected-person risk, but it cannot determine local lawfulness, rights, or acceptable use.
- 07
National Institute of Standards and Technology
Cybersecurity Framework 2.0The 2024 framework supplies a non-prescriptive taxonomy for governing and communicating cybersecurity outcomes. It can structure system and supplier questions, but it does not prescribe controls or prove that an AI project is secure.
- 08
International Organization for Standardization
ISO/IEC 42001:2023 AI management systemsThe public record describes requirements for an organizational AI management system and continuing improvement. The complete standard is paid material, and its record does not establish certification, conformity, or effective operation for any organization.
- 09
International Organization for Standardization
ISO/IEC 23894:2023 AI risk management guidanceThe public record describes customizable guidance for integrating AI-specific risk management into organizational activity. The complete standard is paid material and does not determine local risk acceptance or compliance.
- 10
International Organization for Standardization
ISO/IEC 5259-1:2024 data quality overviewThe public record covers terminology and examples for data quality in analytics and machine learning across the data lifecycle. It is the overview part of a paid series, not a complete data-quality method or proof that local inputs are fit for purpose.
- 11
International Organization for Standardization
ISO/IEC 22989:2022 AI concepts and terminologyThe public record identifies a standard vocabulary for AI concepts. It helps keep model, system, inference, and lifecycle discussions precise, but terminology alone does not choose a use case or validate an implementation.
- 12
Organisation for Economic Co-operation and Development
Explanatory memorandum on the updated definition of an AI systemThe 2024 memorandum explains the OECD definition adopted for its AI Recommendation. This guide uses its inference-centered boundary to distinguish AI from ordinary deterministic automation; other laws and standards can use different definitions.
- 13
European Commission
AI Act regulatory frameworkThe current European Commission page summarizes the EU's risk-based rules and staged application. Duties depend on jurisdiction, role, system, use, and date, and current legislative changes matter. This source is context, not legal advice.
- 14
Information Commissioner's Office
AI and data protection risk toolkitThe UK regulator's toolkit addresses risks to individual rights and freedoms. Its page currently says the guidance is under review because of the Data (Use and Access) Act, so it must not be treated as a settled or universal compliance checklist.
- 15
U.S. Government Accountability Office
Artificial Intelligence Accountability FrameworkThe 2021 federal framework groups accountability practices under governance, data, performance, and monitoring and includes questions for managers and assessors. Its government context can inform inquiry but does not create commercial assurance.
- 16
OWASP Gen AI Security Project
Top 10 for LLM Applications 2025The community project catalogues prominent security risks for applications using large language models. It is a security-awareness resource, not a complete threat model, control baseline, penetration test, or proof of secure operation.
- 17
Microsoft HAX Toolkit
Guidelines for Human-AI InteractionThe provider toolkit presents research-based interaction guidance for initial use, ongoing interaction, error, and change over time. It is useful design input, not a universal interface recipe or evidence of local usability, accessibility, or safe reliance.
- 18
UK National Cyber Security Centre
Guidelines for secure AI system developmentThe 2023 guidance addresses secure design, development, deployment, operation, and maintenance for AI system providers, including those using hosted models and external APIs. Local threats, duties, architecture, and control evidence still need assessment.
- 19
Google Cloud DORA
Test automationThe current delivery guidance emphasizes fast, reliable feedback, developer ownership, continuous testing, and continued manual and exploratory work. Its software-delivery research is an analogy for evaluation discipline, not evidence of an AI business outcome.
- 20
World Wide Web Consortium
Web Content Accessibility Guidelines 2.2The W3C Recommendation supplies testable web-content accessibility criteria for human-facing interfaces. Conformance applies to complete page variations and still requires appropriate evaluation; citing the standard does not establish product accessibility.
Start with one real workflow
A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.
Show Us the WorkflowStart with the free automation readiness checklistOBSERVEQUANTIFYDECIDEBUILD
