AI model selection: costs, control and product continuity.
A model choice needs testing against real work. Data, application workflows, costs and migration options matter alongside benchmark results.
Start with the task
Which is the least expensive model that reliably meets this task's requirements?
I first establish minimum quality, response time, volumes, privacy and continuity. Choosing a provider comes later. A model that excels overall may be unsuitable for a particular activity; a smaller one may be sufficient, if representative tests demonstrate it.
The organisational and economic assessment comes before this choice. Achieving technical answer quality requires separate development and verification work.
Available weights and proprietary services
Open-weight means that the model's weights are available. It does not automatically imply an open-source licence, unrestricted use or lower costs. An open-weight model can also be accessed through an external service.
Epoch AI's May 2026 analysis finds that, since January of that year, leading open-weight models lagged the closed frontier by an average of about four months on its ECI index. This describes a period and a chosen measure, rather than equivalence on every task.
Anthropic's September 2026 cyber comparison provides another narrow example: across 41 V8 vulnerabilities, GLM-5.3 built working exploits in about 12% of attempts, compared with Mythos Preview's 14%. This is not a comparison of general model quality.
My architectural conclusion is to make migration possible when a model's advantage changes, sizing that investment to the project's risks.
The cost of the complete process
Token prices are one part of the bill. Data preparation, infrastructure, retries, human checks, monitoring and maintenance also count. Self-hosting requires GPUs and operational expertise; an API may be more economical for low or irregular workloads.
In McKinsey's May 2026 survey, 93% of the 75 qualified respondents reported exceeding their AI budget. The sample does not represent all companies, but draws attention to the costs of adoption at scale.
A workflow can route different activities to different models: more capability where the problem requires it, cheaper or specialised models for clearly defined steps. Each route must respect data constraints and pass the same quality checks.
How long does it need to keep working?
With weights, runtime and hardware under your control, a configuration can be retained while it meets requirements. A new model release does not require a migration; security updates and component maintenance remain necessary.
An external API has a different lifecycle: OpenAI and Anthropic document deprecations and retirements. Distinguish a mandatory migration, because a service disappears, from a beneficial one, because an alternative improves quality, costs or performance.
I therefore budget for reassessment and plan for continuity. Model independence is an investment: its value grows with the probability and cost of needing a replacement.
Isolate what may change
I would organise the system into parts with distinct responsibilities:
- 01Application and domain
Product rules, permissions and data that remain the application's responsibility.
- 02Workflow and tools
Steps, checks, information retrieval and permitted actions.
- 03Model adapter
Provider-specific formats, instructions, structured outputs and tool calls.
- 04Model and execution environment
The external service or a model hosted on controlled resources.
This is a design proposal, not a universal requirement. Adapters reduce dependencies, but models differ in behaviour: migration often requires changes to prompts, tools and context handling.
A migration must pass the tests
The ability to change models becomes credible with a stable evaluation suite. I compare the current model and the candidate on real requests and edge cases, retaining explicit acceptance criteria.
- Correctness, grounding and regressions against the required behaviour.
- Success of the complete task and of tool calls.
- Latency, concurrency, cost per completed operation and resource consumption.
- Errors, refusals, recovery from failures and compliance with security constraints.
Repeat the measurements after changes to models, data or workflows. A public benchmark helps shortlist candidates; it does not replace application testing.
Assess the choice to wait, too
In my judgement, non-adoption also merits assessment: a competitor might automate an activity or develop expertise before us. This is a strategic scenario to examine, not a prediction that applies to every business.
I propose active caution: experiment within a limited scope, measure value and costs, maintain skills and decide on your own evidence. A project can be postponed or rejected when the benefit does not justify the investment and risks.
Supporting sources
These sources support the facts and evaluation criteria. The architecture and the comparison between adoption and waiting are my methodological proposals, rather than findings demonstrated for every project.
- Epoch AI — open/closed comparison (opens in a new tab): an aggregate index and observation period, not a comparison of a particular use case.
- Anthropic — GLM-5.3 cyber capabilities (opens in a new tab): specific exploit tests, separate from general model quality.
- McKinsey — enterprise AI costs (opens in a new tab): survey and criteria for model selection and task routing.
- OpenAI — API deprecations (opens in a new tab): announcements of service retirement and migration.
- Anthropic — model lifecycle (opens in a new tab): status and retirement of models available through APIs.
- Microsoft — RAG evaluations (opens in a new tab): examples of measures for information retrieval and answer quality.