Skip to main content

Decision guide / LLM providers

Compare the model, the delivery path, and the data boundary.

A provider name is not a complete technical choice. The exact model, endpoint, hosting arrangement, tools, and data terms determine what you can build and what you must operate. Use this dated comparison to form a shortlist, then test it against your own work.

The selection unit

Choose a configuration you can explain and reproduce.

Compare OpenAI, Anthropic, Google, Mistral, and Cohere as possible suppliers of model capability. Then name the actual model version, serving platform, tools, region, and data settings. Reject configurations that fail a mandatory requirement before ranking the remaining options on measured task quality, response time, total cost, and operating effort.

Provider is not model
One catalog can contain different modalities, licenses, release states, and retention rules. A claim about one model does not automatically apply to its siblings.
Model is not platform
A hosted tool, retrieval store, cloud endpoint, and local model server create different dependencies. Record which component supplies each capability and where its data goes.
Specification is not acceptance
A context limit or benchmark score does not establish reliable behavior on your documents, languages, tool contracts, or exception cases. Test the full application path.

Five provider profiles

Start with the difference that matters to your application.

This selection covers general-purpose hosted models and options with downloadable weights. It is not a claim that these are the five largest or best providers. Model examples reflect the official pages reviewed on 10 September 2026. The proposed fit and test questions are Werkon's analysis, not independent benchmark results.

OpenAI

01

Shortlist question: Does a model plus the Responses tool ecosystem fit the workflow?

Reason to evaluate
You want to evaluate a hosted model alongside provider tools such as web search, file search, and function calling.
Documented examples
The current catalog features GPT-6 Astra and GPT-5.6 Sol, Terra, and Luna. It also lists separate audio, image, embedding, and open-weight models.
Integration boundary
Responses supports built-in tools and custom functions; availability depends on the selected model and tool. Remote services add their own data boundary.
Buyer responsibility
Your application still owns permissions, validation, approvals, and evidence that an action completed.
Test to run
Run representative tasks with the actual tool set, reasoning settings, state handling, and error paths you intend to use.
Important limit
Data controls have endpoint, feature, and customer/model exceptions. Zero Data Retention is not a blanket statement that every feature stores nothing.
Change to watch
Review model eligibility, retained application state, tool support, and retirement notices before changing the configuration.

Anthropic

02

Shortlist question: Which Claude model meets the task and the required retention policy?

Reason to evaluate
You want to compare text and image reasoning with tool use across Claude's current model choices.
Documented examples
The overview lists Claude Fable 5.1, Opus 5, Sonnet 5, and Haiku 4.5 with text/image input and text output. Platform availability is model-specific.
Integration boundary
Choose the exact Claude model and serving platform, then verify the features and model identifiers on that path.
Buyer responsibility
Keep action authorization and acceptance outside the model. Review the policy attached to the chosen model, not just a general enterprise description.
Test to run
Compare task completion, unsupported answers, tool selection, and latency at the intended thinking or effort setting.
Important limit
Fable 5.1 is a Covered Model. The standard policy includes minimum retention; a limited eligible ZDR transition arrangement is described separately and is not universal access.
Change to watch
Recheck Covered Model designation, the actual workspace agreement, platform support, and any transitional arrangement.

Google

03

Shortlist question: Which Gemini modality and release state does the application actually need?

Reason to evaluate
You need to assess a catalog spanning text and multimodal tasks, live interactions, and specialized media models.
Documented examples
The fetched Gemini API catalog lists Gemini 3.8 Flash as stable, alongside other Flash models; Gemini 3.1 Pro is listed as preview. Separate live, speech, image, and video models have their own entries.
Integration boundary
Select a specific model endpoint. Stable, preview, latest, and experimental names have different change expectations; do not treat the catalog as one interchangeable API capability.
Buyer responsibility
Verify the actual product and billing arrangement. Gemini API terms do not automatically describe a separate Google Cloud deployment.
Test to run
Test the required input/output modality, tool behavior, long-input accuracy, and release constraints with representative files.
Important limit
Gemini API data terms distinguish paid and unpaid services, with specific EEA, Swiss, and UK provisions. Paid-service treatment does not mean zero retention.
Change to watch
Monitor preview retirement, moving aliases, product-specific terms, and any added grounding or media service.

Mistral

04

Shortlist question: Does the exact model's deployment and license fit your operating plan?

Reason to evaluate
You want to compare hosted use with a model whose weights can be operated under its stated license.
Documented examples
Mistral Medium 3.5 is documented as a multimodal model with open weights under a Modified MIT license. The catalog separately labels Small 4 and Large 3 as Apache 2.0.
Integration boundary
Medium 3.5 lists structured outputs and function calling, while built-in tools are associated with specific agent/conversation endpoints. Match features to the serving path.
Buyer responsibility
Review the selected license and service terms, and assign infrastructure, security, patching, and support ownership for any deployment you operate.
Test to run
Run the same task contract on the intended runtime, including structured output failures, concurrency, and recovery.
Important limit
Open weights do not imply one common license, hosted feature parity, or low operating cost. The catalog also includes third-party models.
Change to watch
Track the actual model publisher, weight revision, runtime, license, and hosted endpoint lifecycle separately.

Cohere

05

Shortlist question: Does Command A+ fit the language, retrieval, and deployment needs?

Reason to evaluate
You want to evaluate text/image input, tool use, citations, and structured output with a separately assessed serving arrangement.
Documented examples
Command A+ uses model ID command-a-plus-05-2026. Its page documents text/image input, text output, reasoning, tool use, citations, and an Apache 2.0 weight release.
Integration boundary
The model page identifies Model Vault for production use and distinguishes Cohere deployments from open-source deployments. Confirm the supported path and terms for your application.
Buyer responsibility
Validate cited claims against the underlying records and keep retrieval permissions and action approval in the application.
Test to run
Test answer support, language-specific errors, tool arguments, and the actual deployment's response time under representative demand.
Important limit
Downloadable weights and a hosted model are different operating choices. Published hardware examples do not prove capacity or economics for your workload.
Change to watch
Recheck model availability, deployment support, data controls, and the retrieval components used with the generator.

A compact comparison

Write down the option you would actually buy or operate.

These are distinguishing checks, not provider scores. A provider can supply more than one viable configuration. Keep contractual requirements as pass/fail conditions rather than points that a higher quality score can offset.

ProviderModel examplesCapability boundaryOperating choiceBefore selection
OpenAIGPT-6 Astra; GPT-5.6 familyModel plus selected Responses toolsIdentify hosted state and every external toolConfirm exact model, feature, and data-control eligibility
AnthropicFable 5.1; Opus 5; Sonnet 5; Haiku 4.5Claude model and platform feature supportMatch workspace settings to model policyResolve Covered Model retention requirements
GoogleGemini 3.8 Flash; 3.1 Pro previewModality, endpoint, and release stateDistinguish Gemini API from other productsConfirm product terms and preview/change exposure
MistralMedium 3.5; Small 4; Large 3License and endpoint differ by modelChoose hosted or operated weights explicitlyConfirm license, runtime features, and support owner
CohereCommand A+Generation, citations, tools, and serving pathAssess Model Vault or weight deploymentValidate retrieval evidence and deployment conditions

From catalog to decision

Compare accepted work under the same conditions.

A useful selection produces a reproducible record: requirements, candidate configurations, evaluation cases, results, exclusions, and a named decision owner. Do not send sensitive evaluation data until the proposed processing boundary has been approved.

  1. 01

    Set mandatory constraints

    Name the task, allowed data, required modalities, action permissions, deployment restrictions, and unacceptable failures. Exclude configurations that cannot satisfy them.

  2. 02

    Freeze each candidate

    Record model ID, version, endpoint, region, tools, prompt, retrieval setup, reasoning settings, and release state. A model name alone is not a reproducible experiment.

  3. 03

    Test representative cases

    Use normal work, difficult inputs, missing evidence, conflicting records, refusals, and recovery cases. Score correctness and task completion against criteria set before seeing the answers.

  4. 04

    Measure the full cost

    Count model and tool usage, retries, retrieval, review, hosting, monitoring, and maintenance. Compare latency distributions and cost per accepted task at the required quality, not headline token prices alone.

  5. 05

    Choose and rehearse change

    Record why the configuration passed, where it remains weak, and when it must be reconsidered. Exercise a fallback or migration with the same data and permission constraints.

Keep the comparison current

A reviewed date needs a change trigger behind it.

This article is a documentation snapshot, not a live status service. Refresh the affected claims when a model, service, policy, license, or deployment changes, and before using the comparison for a purchase or production decision.

Model and service identity
Track aliases, snapshots, previews, retirement dates, and tool versions. Re-run affected evaluations when behavior can change, even if the provider name stays the same.
Data and contractual scope
Separate training use, abuse monitoring, application state, processing location, and third-party tools. Read exceptions and account eligibility instead of turning a marketing label into an absolute.
Failure and fallback behavior
A fallback provider must independently pass the same data and task requirements. Decide whether an outage should queue work, reduce scope, request review, or stop.
Evidence ownership
Keep evaluation inputs, rubrics, outputs, traces, configuration, and decision notes in systems your team can access. A dashboard screenshot alone is not enough to reproduce the choice.

Questions behind the shortlist

Resolve these before declaring a winner.

Werkon integrates AI into business systems. The useful recommendation is a tested configuration for a defined workload, with clear limits and ownership.

Which provider is best?
This comparison does not establish a universal winner. Form a shortlist from mandatory requirements, then compare actual configurations using the same acceptance cases. Different parts of one application may justify different models.
Is a larger context window always better?
It permits a larger input envelope, but does not establish that the model will find, reconcile, and cite the right evidence. Test relevant facts at different positions, contradictory records, and realistic document lengths.
Are downloadable weights automatically cheaper or more private?
No. Your deployment determines access, logs, network paths, updates, and operating cost. Include infrastructure utilization, engineering effort, security work, and license obligations in the comparison.
Can we switch providers through a common API?
A common request shape can reduce integration work, but it does not prove equivalent tool semantics, output quality, safety behavior, state handling, or data terms. Treat a switch as an application change with its own acceptance evidence.
Why are there no benchmark rankings or price winners here?
We have not run an independent paid comparison for your workload. Published specifications help define candidates; cost and quality need the actual task, settings, demand, failure rate, and review effort. Record current price inputs when you run that comparison.

Source basis

Sources behind the control model.

  • 01

    OpenAI

    API model catalog

    Current model examples and modality families; provider descriptions are not independent performance findings.

  • 02

    OpenAI

    Using tools

    Responses tools, custom functions, and remote services have distinct integration boundaries.

  • 03

    OpenAI

    Data controls in the OpenAI platform

    Retention, application state, feature eligibility, and model/customer exceptions require configuration-specific review.

  • 04

    Anthropic

    Claude models overview

    Model lineup, supported input/output types, and model-specific platform references.

  • 05

    Anthropic

    Covered Models

    Standard retention policy and the separately described, limited eligible transition arrangement must be read together.

  • 06

    Google

    Gemini API models

    Directly fetched catalog, including stable and preview distinctions; search snippets may lag this page.

  • 07

    Google

    Gemini API Additional Terms of Service

    Product-specific paid/unpaid data treatment and regional provisions; not a substitute for reviewing the applicable agreement.

  • 08

    Mistral AI

    Models overview

    Model and publisher distinctions, including different license labels within the same catalog.

  • 09

    Mistral AI

    Mistral Medium 3.5

    Open-weight release, stated license, and endpoint-specific features for this model.

  • 10

    Cohere

    Command A+

    Model ID, modalities, capabilities, deployment distinctions, and Apache 2.0 release; hardware examples are not workload capacity guarantees.

[ WORKFLOW / SYSTEMS AUDIT ]
THE FIRST ENGAGEMENT

Start with one real workflow

A Systems Audit is the usual starting point. If the opportunity is already clear, we can move directly into a focused build.

Show Us the WorkflowStart with the free automation readiness checklist

OBSERVEQUANTIFYDECIDEBUILD