Open weights don't make enterprise AI cheaper. They make it portable.
Take a hypothetical case. You run a software startup. Your product puts an AI agent to work on customers' business data: pricing models, product catalogues, approval rules. In development it runs on a closed frontier API, and it works.
Then the pipeline fills with the customers you actually want: a bank in Jakarta, an insurer in Singapore, a healthcare group in Germany, a manufacturer in Australia. Each sends a vendor due-diligence questionnaire, and the same four questions show up in every one:
- Which model processes our data, and in which country?
- Will you tell us before that model changes?
- Can our regulator audit your providers?
- What happens if your model provider withdraws the service?
Then each market adds its own twist. The Indonesian bank's regulator expects its systems processed in Indonesia unless approved otherwise. The German healthcare group needs processing in the EU on certified infrastructure. The Australian customer has to notify its regulator before data goes offshore. None of this is the customer being difficult. It is their regulators, and it flows down to you as their vendor.
With one closed API, the honest answers are thin: the provider's current model, wherever the provider runs it; only when the provider tells us; probably not; we move to whatever the provider offers next. And you cannot place that model in Jakarta and Frankfurt at the same time.
An open-weight model changes the answers. You can name the exact version, pin it, run the same weights in each region your customers need, and move them if a host fails you. One product, one set of evals, many jurisdictions. That is what this article is about.
It is not about the leaderboard. Benchmark leads now expire in months: Qwen's 27B beat its own 397B flagship in April and was overtaken by its successor in August. It is not mainly about price either. Closed providers cut prices too, and a choice made only on price gets reversed at the next price cut.
The position: open weights turn the model from a vendor relationship into a component you control. They don't make enterprise AI cheaper. They make it portable. And portability is what regulators in Singapore and the EU are now writing into their rules.
The short version:
- Cost: open vs closed is a cost decision. Hosted open models are 13 to 17 times cheaper per token at the frontier and mid tiers. Hosted vs self-hosted is a control decision; owning GPUs rarely beats renting the same open model.
- Default path: start on a hosted open-weight API. Move a workload in-house when one of three triggers fires: a rule that requires local processing, steady high volume, or fine-tuning.
- Regulators: MAS, IMDA, the EU AI Act and DORA do not prefer open or closed. They ask for models you can test, audit, pin and exit. Open weights make those answers easy.
- This quarter: separate the model choice from the hosting choice, classify workloads by residency rule, build your own eval set, set up a provenance check before the first download, and test the exit swap once a quarter.
In this article:
Part 1: The argument
- What open weights actually change: exit, residency, cost behaviour, customisation, latest vs predictable, deployment control, evals, the interface layer
- What it costs: open is cheaper, owning the GPU often is not: open is 13-17x cheaper than closed, owning the GPU rarely beats renting
- What regulators now ask for: MAS, IMDA, the EU AI Act and DORA on third-party models
- Residency by industry: where inference has to stay: the rules that force local processing, and the ones that only require a local copy
- Who is already doing it: adoption data, named enterprise cases, sovereign AI in APAC
- The model market in three points: licence traps, capability vs provenance, the gap to closed frontier
- Open weights still need diligence: supply-chain attacks, what self-hosting does not fix, pinning vs deprecation
- Self-host or API: the decision: a table for the actual call
- What running it looks like: evidence from a working setup
- Where closed APIs still win: the counterargument, taken seriously
- What to do with this: five decisions for this quarter
Part 2: Reference
- Which models, on what hardware: hardware tiers, 16 model families vendor by vendor, best pick per use case
- Where to rent open weights: 18 hosts compared on models, APAC and EU regions, jurisdiction, data defaults
Part 1: The argument
What open weights actually change
Four things shift when you hold the weights, not just an API key.
| Closed API | Open weights | |
|---|---|---|
| Exit | Model retired or repriced on the provider's schedule | Same weights run on another host, or your own hardware, unchanged |
| Data residency | Wherever the provider's region map allows | Wherever you put the GPU |
| Cost behaviour | Per token, scales linearly with usage | Per GPU-hour when self-hosted, flat once the hardware is busy |
| Customisation | Whatever fine-tuning the provider exposes | LoRA or full fine-tune on your own data, and you keep the result |
Exit is the one most teams underrate. The same open-weight model is served today by several independent inference providers. If one raises prices, degrades latency or drops the model, you move the endpoint. Your prompts, evals and agent code stay the same. No closed model gives you that.
Residency becomes an infrastructure decision, not a contract negotiation. You do not ask a provider whether they can process data in Jakarta. You deploy in Jakarta.
Cost becomes predictable at scale. API spend tracks usage one-to-one, so a successful rollout is also a budget surprise. A self-hosted node costs the same at 30% utilisation as at 90%. That is worse for spiky, low-volume work and much better for steady, high-volume work.
Customisation becomes an asset you own. A LoRA adapter trained on your contracts, tickets or product catalogue is a file on your storage. It does not disappear when a vendor deprecates the base model.
So when you evaluate an open-weight model, ask about portability first. Benchmark rank comes second.

Latest model or predictable model
The instinct is to run the latest and greatest. For some work that is right: a coding assistant or a research agent gets better every time the model does, and nobody minds if it answers differently next week.
Many enterprise use cases want the opposite. An agent that extracts fields from contracts, prices a quote or changes a customer's system configuration has to behave the same way tomorrow as it did on the day it was validated. Every model change means re-running the evals, re-checking the edge cases and, in a regulated customer, telling the regulator's control framework why the output changed. A few benchmark points are not worth that.
| Latest model | Pinned model | |
|---|---|---|
| Typical workloads | Coding assistants, research, drafting, exploration | Extraction, classification, pricing, configuration changes, regulated decisions |
| What matters | Capability ceiling | Same input, same output; validated behaviour |
| Model change | Welcome | A release with regression testing and sign-off |
| Who sets the upgrade date | Whoever ships the newest model | You, after your evals pass |
Closed APIs serve the left column well. The right column is where they struggle: the provider retires versions on its schedule, and behaviour can shift under an unchanged model name. Open weights let you run both columns at once. Use the newest model where it pays, and keep a pinned, validated version for the workloads that have to stay predictable, upgrading them when you decide.
Deployment control is the point, not the licence
Look at what the largest Mistral enterprise deals actually buy. BNP Paribas, CMA CGM and the French armed forces run Mistral's commercial models on their own or French infrastructure, not necessarily the Apache-licensed weights. What they pay for is the right to run the model where they decide, under their own controls. The open licence is one way to get that right. A commercial on-premises licence is another. Either way, the sovereignty comes from deployment control, and that is what your regulated customers are asking about.

Your evals are the asset you carry
Models change every quarter. Hosts come and go. What survives every change is your evaluation set: the few hundred real tickets, contracts, quotes or configuration tasks with known correct outcomes, scored the same way every time.
That set is what makes the rest of this article work. It tells you whether an open model is good enough to replace a closed one. It tells you whether a new version can replace the pinned one. It gives a regulator evidence that behaviour did not change after a host switch. Building a demo on a frontier model is easy. Proving a production workload still behaves after the model underneath changes is the hard part, and only your own evals do that.
Treat the eval set as a product asset with an owner, version control and a release gate. It is the one piece of the AI stack you fully own, whichever model you run.
The interface layer makes the swap cheap
Portability only pays if switching models is cheap. Two things make it cheap today. Most hosts, open and closed, expose an OpenAI-compatible API, and many also accept the Anthropic API format. And agent tooling increasingly connects through the Model Context Protocol (MCP), so tools and data sources are wired to the agent, not to a specific model.
Build on those two layers and changing the model under an agent is a configuration change plus an eval run. Build on one provider's proprietary features (its own agent runtime, assistants API or fine-tuning format) and every switch becomes a migration project. The portability argument is decided in your architecture long before it reaches a procurement questionnaire.

What it costs: open is cheaper, owning the GPU often is not
The title says open weights do not make AI cheaper. That needs precision, because two different comparisons get mixed up.
Open model vs closed model: open wins on price. List prices per million tokens, fetched October 3, 2026:
| Tier | Open-weight, hosted (input / output) | Closed (input / output) | Output price gap |
|---|---|---|---|
| Frontier | Mistral Large 3: $0.50 / $1.50. DeepSeek V4 Pro: $1.30 / $2.60. GLM-5.3: $1.40 / $4.40 | Claude Opus 5.5: $4 / $20. GPT-6-astra: $10 / $50 | ~13x (Large 3 vs Opus 5.5) |
| Mid | gpt-oss-120b: $0.15 / $0.60. DeepSeek V4.1 Flash: $0.30 / $1.20 | Claude Sonnet 5.5: $2 / $10. GPT-6.1-sol: $2 / $10 | ~17x (gpt-oss-120b vs Sonnet 5.5) |
| Small | gpt-oss-20b: $0.075 / $0.30. Qwen3-30B-A3B (Alibaba, Singapore region): $0.20 / $0.80 | Claude Haiku 4.5: $1 / $5. GPT-6-luna: $0.10 / $0.50 | ~17x vs Haiku 4.5; GPT-6-luna is priced like an open model |
At the small end the gap is closing: OpenAI prices GPT-6-luna at $0.10 / $0.50, in open-model territory. Same-tier comparisons also flatter open models, because closed models still lead on quality. Mozilla's 2026 open-source AI report puts the gap at roughly 6x at comparable quality. Still large.
Self-hosting vs renting the same open model: renting usually wins. Hosted open-weight providers run at high utilisation and batch across customers. You will not beat that on a lightly used GPU.
A worked estimate, with every assumption visible:
| Assumption | Value |
|---|---|
| Model | gpt-oss-120b on one H100, vLLM, ~2,000 output tokens/s aggregate (±50%) |
| Rented H100 | $3.99/hour list price, ~$2,900/month |
| Owned 8x H100 in a Singapore colo | ~$15,600/month: server amortised over 36 months, power, rack, half an engineer |
| Input:output ratio | 3:1 |
| Scenario | Cost per million output tokens |
|---|---|
| Rented H100 at 100% utilisation | $0.55 |
| Rented H100 at 50% | $1.11 |
| Rented H100 at 25% | $2.22 |
| Hosted gpt-oss-120b API | $0.60 |
| Claude Sonnet 5.5 API | $10.00 |
A rented GPU only matches the hosted open API at near-full utilisation. The owned 8-GPU node breaks even against the hosted open API at roughly 35% average utilisation, and at about 70% if real throughput is half the assumption. Published analyses land in the same place: a Spheron study of Mistral Large 3 on 8x H200 concluded the cluster could not reach break-even volume against the API, and justified self-hosting on residency, not cost.
So the economics split cleanly:
- Open vs closed is a cost decision. If an open model is good enough for the task, the hosted version is an order of magnitude cheaper.
- Hosted vs self-hosted is a control decision. You pay for it unless your volume is steady enough to keep GPUs busy, or a residency rule leaves no choice.
That is the portability argument in numbers. Start hosted, and move only when control is worth the premium.

What regulators now ask for
None of the new rules prefers open or closed models. What they prefer is a model you can test, audit, pin and leave. Read them closely and that is a list of open-weight properties.
| Instrument | Status (Oct 2026) | What it asks of a firm using third-party AI |
|---|---|---|
| MAS Guidelines on AI Risk Management (Singapore, all financial institutions) | Consultation Nov 2025; MAS said in Aug 2026 the final text is coming "soon", with a proposed 12-month transition | AI inventory including third-party models. Test the vendor model on your own use case and data. If the vendor discloses too little, limit use under §4.3 or test compensatorily under §4.11. For open-source models, check provenance and training-data integrity (§4.11(c)). Assess concentration on key providers, plan for a vendor dropping support, secure audit rights and change notifications (§4.11(d)-(f)) |
| IMDA Model AI Governance Framework for Agentic AI (Singapore) | v1.0 Jan 2026, v1.5 May 2026; voluntary | Limited visibility and control over third-party components is a named risk factor. Require disclosures, scoped credentials and tool-call logging, or reassess the deployment |
| IMDA Starter Kit for Testing LLM Applications (Singapore) | Jan 2026; voluntary | Pre-deployment testing for hallucination, bias, harmful content, data leakage and adversarial prompts. Warns explicitly about backdoored fine-tunes from unverified sources |
| EU AI Act, amended by the Digital Omnibus | GPAI duties live since Aug 2025, Commission enforcement since Aug 2026; high-risk obligations moved to Dec 2027 | Deployer duties apply whether the model is open or closed. The open-source exemption (Art. 53(2)) only relieves the model's publisher, never your deployment, and never models above 10^25 FLOP |
| EU Commission GPAI guidelines | Jul 2025 | You become the "provider" of a modified model only if the modification uses more than a third of the original training compute. A LoRA fine-tune keeps you a deployer |
| DORA (EU financial sector) | Applies since Jan 2025 | An LLM API is an ICT third-party service: register entry, contract terms on data location, audit and termination, a tested exit strategy |
Three consequences follow.
Opacity has a price, and the deployer pays it. MAS §4.3 and §4.11 are explicit: if your vendor will not tell you enough, you run more tests or you use the model for less. With weights in your own environment, the information gap does not exist.
Silent model updates become a control finding. MAS wants notification of third-party model changes and an impact assessment. Pinned weights do not change unless you change them. With an API, you need version pinning in the contract and regression tests on every release.
Open weights are not a free pass. The same MAS paragraph that makes the exit case also names open-source models with weak security controls as a risk. A Hugging Face download is a software supply-chain dependency. Treat it like one: verified source, checksums, scanning, red-teaming before production.
Residency by industry: where inference has to stay
Most discussions of data residency blur two different rules. "Store locally" means a copy of the data must sit in-country; processing elsewhere can still be allowed. "Process locally" means the compute itself must run in-country. For an LLM, only the second one forces inference into the country.
| Market | Sector | Rule | Effect on an LLM |
|---|---|---|---|
| Singapore | Government | GovTech GenAI control under IM8 | Overseas-hosted models: data up to Restricted only. Singapore-hosted: up to Confidential. Provider must not log, store or train on inputs |
| Singapore | Banks | MAS Notice 658 on outsourcing (in force Dec 2024) | No localisation. Register, due diligence, audit rights for the bank and MAS, exit planning |
| Indonesia | Banks | POJK 11/2022 | Systems placed in Indonesian data centres; offshore placement needs OJK approval. Process locally by default |
| Indonesia | Insurance, finance, fintech | POJK 4/2021 | Same pattern as banks |
| Malaysia | Banks, insurers | BNM RMiT (reissued Nov 2025), Outsourcing policy | Consult BNM before first public-cloud use for critical systems; prior approval for material outsourcing, with processing locations listed |
| Thailand | Banks | BOT FPG 19/2559 | Public cloud for critical IT needs BOT approval 30 days ahead |
| Vietnam | All data controllers | PDP Law 91/2025 (from Jan 2026) | Offshore processing allowed, but a transfer impact assessment must be filed within 60 days |
| India | Payments | RBI payment data directive (2018) | Store locally. Processing abroad allowed, data back in India within 24 hours |
| India | Securities | SEBI cloud framework (2023) | Storage and processing in empanelled Indian data centres. Process locally |
| China | Banking, payments | PBOC rules on personal financial information | Stored, processed and analysed in China. Process locally |
| Korea | Financial | Electronic Financial Supervisory Regulation, network separation | Credit and ID data on cloud must sit in Korea; internet-based LLM services only via sandbox exceptions. Process locally |
| Australia | Banks, insurers, super | APRA CPS 230 (Jul 2025) | Notify APRA before any material offshoring. Regulator access rights, orderly exit |
| Australia | Health | My Health Records Act s77 | No holding or processing outside Australia. Process locally |
| EU | Financial | DORA | Contract must state processing regions; exit strategy; non-EU critical providers need an EU subsidiary |
| Germany | Health | §393 SGB V | Processing in Germany, EU/EEA or adequate country, plus BSI C5 attestation |
| France | Sensitive state data | SREN law + SecNumCloud | Providers must be immune to non-EU extraterritorial law. US hyperscaler APIs excluded |
Read the table by column, not by row, and three patterns show.
Outright bans are rare. Approval and notification gates are common. Malaysia, Indonesia, Thailand, Australia and the Philippines all let a bank use an offshore service, after the regulator has seen it. The real cost of a closed API in APAC banking is lead time and paperwork per use case.
Regulators keep audit rights over your providers. MAS 658, RBI, APRA CPS 230 and DORA all require that the regulator can inspect the service provider. Frontier-model API terms rarely concede that. A model running inside your own environment needs no such negotiation.
Exit plans are mandatory almost everywhere. RBI, APRA, DORA and MAS all require one. From January 2027 the EU Data Act also removes cloud switching charges. The easiest exit plan to evidence is a model you can redeploy elsewhere tomorrow.
For a group operating across five ASEAN markets, this decides whether one AI platform covers the region or whether each country gets its own exception file. A closed model is available where its provider has a region. An open-weight model is available wherever there is a GPU.
Two more regional factors count. Some regional cloud providers serve open-weight models from Singapore endpoints with OpenAI- or Anthropic-compatible APIs, so you get regional latency without changing agent code. And open-weight families with strong Asian-language training can be fine-tuned further on Bahasa, Thai or Vietnamese corpora you own.

Who is already doing it
The adoption numbers point two ways. Menlo Ventures' December 2025 survey of 495 US enterprises found open-source models at 11% of enterprise LLM use, down from 19% a year earlier. Developer usage moved the other way: open-weight models carried about a third of tokens on OpenRouter by late 2025, and Hugging Face counted 151,448 Qwen derivatives by August 2026. Mozilla's 2026 report names the gap: 51% of developers run open models in production, against 63% for closed.
The gap is operations, not model quality. Where enterprises do commit, sovereignty and customisation are the stated reasons.
| Organisation | Model | Deployment | Why |
|---|---|---|---|
| BNP Paribas | Mistral commercial models | On-premises, all business lines (2024, extended May 2026) | EU regulation, data sovereignty |
| CMA CGM | Mistral | €100M over five years, Mistral engineers embedded in Marseille | Customisation |
| French Ministry of Armed Forces | Mistral | Framework agreement 2026-2030, French infrastructure, fine-tuned on defence data | Sovereignty |
| Stellantis | Mistral | In-car assistant, engineering analytics | Customisation |
| Orange | Codestral | Hosted on Orange Business infrastructure, resold to business customers | Sovereignty |
| Deutsche Telekom | SOOFI, ~100B open European LLM | Industrial AI Cloud, Munich, 10,000+ GPUs | Sovereignty |
| GoTo + Indosat (Indonesia) | Sahabat-AI, Gemma-based, 70B version 2025 | Trained and served in-country on GPU Merdeka; runs in GoPay | Local languages, sovereignty |
| SK Telecom (Korea) | A.X 4.0, Qwen2.5 continued pre-training | Open weights, local deployment | Korean token efficiency |
| SCB 10X (Thailand) | Typhoon, on Mistral, Qwen and Gemma bases | Open weights | Thai language |
| Rakuten (Japan) | Rakuten AI 3.0, ~700B MoE on a DeepSeek-V3 architecture | Apache 2.0 | Japanese model, government-subsidised |
Two cases keep the picture honest. OCBC built its internal GPT on Azure OpenAI in a controlled environment: a sovereignty-sensitive bank that chose closed. And the French state's Albert assistant, built on Llama and Mistral on government infrastructure, was not rolled out further "in its current form" in January 2026 after a pilot at 48 sites. Open weights do not rescue a weak product.
Sovereign AI in APAC runs on open weights
| Country | Programme | Open weights? | Base |
|---|---|---|---|
| Singapore | NAIRD, S$1bn+ for 2025-2030; S$70M national LLM programme (SEA-LION, MERaLiON) | Yes | Foreign open bases: Gemma, Llama, Qwen, now Nemotron |
| Indonesia | Sahabat-AI on GPU Merdeka; Danantara sovereign AI infrastructure push (Sept 2026) | Yes | Gemma, Llama |
| Malaysia | YTL ILMU | ILMU 1.0 closed; ILMU-Nemo-30B on NVIDIA Nemotron | Foreign open bases |
| Thailand | ThaiLLM (NSTDA), Typhoon | Typhoon yes | Qwen, Gemma, Mistral |
| Japan | GENIAC subsidies, ABCI 3.0 | Mixed | Rakuten on DeepSeek-V3 |
| Korea | Independent AI Foundation Model Project | Yes (K-EXAONE 2.0, A.X K2 under Apache 2.0) | From scratch by rule: Naver was cut in January 2026 for using foreign weights |
| India | IndiaAI Mission, ~US$1.1bn | Yes (Sarvam 30B and 105B, Apache 2.0) | From scratch |
The pattern: Southeast Asia builds sovereign models by continued pre-training on foreign open weights, increasingly Chinese ones. Korea and India train from scratch and still release open weights. Either way, sovereignty means where the model runs and who can adapt it. None of these programmes would exist on closed APIs.

The model market in three points
The model-by-model detail is in Part 2. Three points from it shape the argument.
"Open weights" does not mean "open licence". Apache 2.0 and MIT cover most of the list. The exceptions bite exactly where enterprises sit. Mistral Medium 3.5 and Devstral 2 grant no rights to companies with more than $20M monthly revenue, unless they buy a commercial licence. Qwen's largest models and Kimi K3 need a separate agreement if you resell model access at scale. MiniMax M3 bans military use. Llama 4 withholds multimodal rights from EU-domiciled companies. The NVIDIA Open Model Licence terminates automatically if you bypass its guardrails without an equivalent replacement, which matters the day you fine-tune refusals away. And licences change between versions of the same family: Gemma 4 moved to Apache 2.0, GLM-5.3 added a clause GLM-5 did not have. Review per checkpoint, not per family. Read the LICENSE file, not the model card headline.
Capability and provenance pull in opposite directions. On vendor-reported agentic coding benchmarks, the top open tier is Chinese: DeepSeek, Kimi, GLM, MiMo. The strongest non-Chinese open models (Nemotron 3 Ultra, Inkling, Mistral Large 3) sit roughly a tier below. The most transparent models are Western: Nemotron 3 Ultra publishes its training data, Apertus and Olmo publish data, code and recipes. If procurement rules out Chinese origin, you pay for it in capability. Self-hosting removes the data flow to the vendor's country. It does not change what the model was trained to do. That part is evaluation work.
The gap to closed frontier is real but narrow. On the independent Artificial Analysis Intelligence Index, the best open model scores 46 against 58 for the leading closed model and 53 for the rest of the frontier (October 2026). The gap is widest on long-horizon terminal tasks and factual recall, and narrowest on single-session coding. In practice, top open models match the closed frontier of roughly one generation ago.
For most enterprise workloads (extraction, classification, RAG over internal documents, coding agents on a known codebase), one generation behind on a model you control beats the frontier on a model you rent. Does your use case actually need the last 10 points?

Open weights still need diligence
Holding the weights moves risk, it does not remove it. Three risks move onto your side of the line: the files you download, the behaviour baked into them, and the upgrades nobody forces on you.
The files: a software supply chain
Model hubs are package registries, and they get attacked like package registries.
| Date | Incident | What happened |
|---|---|---|
| Mar 2024 | JFrog finds ~100 malicious models on Hugging Face | Pickle payloads opened a reverse shell the moment the model loaded |
| Feb 2025 | ReversingLabs "nullifAI" | 7z-compressed PyTorch files slipped past both the loader checks and Picklescan, and still executed |
| Apr 2025 | CVE-2025-32434 (CVSS 9.3) | torch.load(weights_only=True), the standard "safe" setting, could still execute code on PyTorch 2.5.1 and earlier |
| 2025 | Three PickleScan zero-days (CVSS 9.3 each) | Scanner bypasses, fixed in PickleScan 0.0.31 |
| May 2026 | HiddenLayer: fake "privacy-filter" repo | Impersonated an OpenAI release, hit #1 trending with ~244k downloads in 18 hours. The weights were clean; the bundled loader.py dropped an infostealer |
The last case is the lesson. Scanners check the weights. The attack came through the code shipped next to them.
| Check before the first load | Why |
|---|---|
| Safetensors or GGUF only; reject pickle formats | No executable code at load time |
trust_remote_code=False; review any bundled loader code line by line |
The 2026 attack vector |
| Official publisher repo, pinned by commit hash | Download counts and likes can be faked |
| Verify OpenSSF Model Signing signatures where published (NVIDIA NGC signs every model) | Proves file integrity and publisher, not safe behaviour |
| Two scanners, PyTorch 2.6 or later | Single scanners have missed real payloads |
| First load in a sandbox without network access | Contains anything the scanners missed |
| Serve from an internal mirror, record an AI-BOM (CycloneDX or SPDX 3.0) | The inventory MAS asks for, with version and licence attached |

The behaviour: what self-hosting does not fix
Origin concerns about Chinese models are mostly about data flowing to servers under PRC jurisdiction. Australia, the Czech Republic, Italy, South Korea, Taiwan and several US states acted against DeepSeek's app and service in 2025. Self-hosting solves that part. It does not solve what is inside the weights.
| Concern | Solved by self-hosting? | Evidence |
|---|---|---|
| Data sent to the vendor's country | Yes, with no outbound network access | |
| App telemetry, vendor privacy policy | Yes | |
| Censorship or narrative bias | No | R1dacted study (May 2025): DeepSeek R1 censorship persists when run locally. NIST CAISI (Dec 2025): Kimi K2 Thinking heavily censored in Chinese, much less in English |
| Weak resistance to jailbreaks and agent hijacking | No | NIST CAISI (Sep 2025): DeepSeek R1-0528 agents 12x more likely to follow malicious instructions than US frontier models |
| Context-triggered weaker code | No | CrowdStrike (Nov 2025): politically sensitive context in prompts raised vulnerable-code rates in DeepSeek R1 by up to ~50% |
| Planted backdoors | No, and detection is immature | Anthropic, UK AISI and the Alan Turing Institute (Oct 2025): ~250 poisoned documents enough to backdoor models up to 13B |
| Procurement bans | Read the wording | Australia and the Czech Republic ban "products", which can cover self-hosted weights |
Two caveats keep this honest. CAISI is a US government body comparing against US models. And the same CAISI work found DeepSeek V4 Pro roughly eight months behind the US frontier and more cost-efficient than GPT-5.4 mini on most of its benchmarks. Origin is a risk to test, not a verdict.
The practical answer is the same for every origin: put a guard model in front (Llama Guard 4, Qwen3Guard for non-English, IBM Granite Guardian 4.1 for RAG groundedness), and run your own red-team set in every language your users write in.
The upgrades: pinning cuts both ways
Closed APIs retire models on the provider's schedule.
| Provider | Stated minimum notice | Example |
|---|---|---|
| OpenAI | 6 months for GA models, as little as 2 weeks for previews | gpt-4.5-preview: announced April 14, shut down July 14, 2025 |
| Anthropic | 60 days for public models | Claude Haiku 3.5: notice December 19, 2025, retired February 19, 2026 |
| Google Gemini | Listed dates are "earliest possible" | gemini-2.0-flash: about 16 months from release to shutdown |
Behaviour also changes behind an unchanged model name. OpenAI rolled back a sycophantic GPT-4o update in April 2025. Anthropic's September 2025 postmortem traced degraded answers to three infrastructure bugs, at one point hitting about 16% of requests on one model.
Pinned weights do not change unless you change them. That is the control MAS asks for. The cost is that patches, security fixes and upgrades are now your release process, not your vendor's.

Self-host or API: the decision
Open weights do not mean self-hosted. The same weights are available three ways: a hosted open-weight API, a dedicated deployment in your cloud tenancy, or your own hardware. The decision is per workload, not per company.
| Situation | Hosted open-weight API | Dedicated in your VPC | Own hardware |
|---|---|---|---|
| Exploring, spiky or low volume | Best fit | Overkill | Overkill |
| "Store locally" rule (copy in-country) | Fine with local storage design | Good fit | Good fit |
| "Process locally" rule (inference in-country) | Only with an in-country region | Good fit in a local region | Best fit |
| Regulator needs audit access to the provider | Hard to negotiate | Good fit | Best fit |
| Steady high volume, predictable load | Linear cost | Good fit | Best fit |
| Domain fine-tune needed | If provider serves adapters | Good fit | Best fit |
| No MLOps capacity | Best fit | Needs a platform team | Needs a platform team |
| Air-gapped or classified | Not possible | Not possible | Only option |
The pattern that works: start on a hosted open-weight API, build your evals against it, and move the workload in-house when one of three triggers fires. A process-locally rule, steady volume, or fine-tuning. Because the weights are identical, the move does not invalidate your evals.
With a closed model, that path does not exist. You can only renegotiate the contract.

What running it looks like
This is not theory. The rAInvent stack runs open-weight models daily, in three ways:
- Self-hosted on own hardware: open-weight models served locally for work that should not leave the machine.
- Hosted open-weight APIs: heavy use through Fireworks AI, Alibaba Cloud and other open-weight inference platforms.
- Mixed agent workers: coding, review and test agents routed across DeepSeek, Kimi, gpt-oss and MiniMax, alongside Claude and Gemini. Each role gets the model that fits its cost and quality bar.
Routing like this is only possible because most of the workers are open-weight models served by more than one provider. When one provider is slow or one model underperforms on a task, the change is one line of config.
The pattern is older than LLMs. The In Mind Cloud CPQ platform was built entirely on open source, including contributions back to the Pellet OWL reasoner. The reasoning layer of a commercial product sat on a component the team could read, fix and ship without waiting for a vendor. Open-weight models bring that property to the model layer.
Where closed APIs still win
The counterargument deserves a straight answer.
- Frontier reasoning: the strongest closed models still lead on the hardest multi-step reasoning and long-horizon agent tasks. If that gap is your use case, pay for it.
- Safety and alignment are now yours: fine-tune aggressively and you own the validation of the resulting model. API providers carry that work for you.
- Supply-chain diligence is now yours too: provenance, integrity checks, guard models and red-teaming are your job, and MAS names it explicitly.
- Confidential computing is narrowing the privacy gap: NVIDIA H100 and Blackwell GPUs can run inference inside attested enclaves, with roughly 1-3% overhead on Blackwell when configured correctly. A provider that can prove it cannot read your prompts answers part of the residency objection. It does not answer exit, pinning or audit.
- Operations are real work: GPU capacity planning, inference servers like vLLM, monitoring, upgrades. Without a platform team, self-hosting turns a model problem into an infrastructure problem.
- Utilisation risk: a GPU node idling at 15% is more expensive per token than any API.
None of these argue against open weights. They argue against self-hosting everything. The hosted open-weight API covers most of them while keeping the exit open.

What to do with this
- Separate the model choice from the hosting choice. Write them down as two decisions with two owners.
- Classify each AI workload against the residency table. Store-locally, process-locally, or neither. Only the second forces local inference.
- Run your evals on at least one open-weight model now, on your own tickets, documents or contracts, not public benchmarks. You need that baseline before a regulator or a residency rule forces the move.
- Build the provenance check before the first download. Approved sources, checksums, licence review, a red-team pass. MAS §4.11(c) will ask for it.
- Write the exit plan as a deployment, not a document. If your agent stack speaks an OpenAI- or Anthropic-compatible API, swapping in an open-weight model is a config change. Test that swap once a quarter and you have the exit evidence DORA, APRA and RBI ask for.
The model you pick this quarter will be outdated in two. Whether you can still run it, move it and adapt it after that depends on the decision you make now.
Part 2: Reference
The material behind the argument, for when you shortlist models and hosts.
Which models, on what hardware
Self-hosting starts with one question: what fits on hardware you can actually buy? The answer moved a long way in 2026. A single workstation GPU now runs a model that scores within reach of last year's frontier on agentic coding.
| Tier | Hardware | Memory | What fits (4-bit weights) | Examples |
|---|---|---|---|---|
| Workstation GPU | RTX 4090 / 5090 | 24-32 GB | ~30B dense, ~35B MoE | Qwen3.8-27B, Gemma 4 31B, Granite 4.2 30B, gpt-oss-20b |
| 128 GB unified memory | DGX Spark, AMD Strix Halo, Mac 128 GB | 128 GB | ~120B MoE | gpt-oss-120b, Mistral Small 4, Nemotron 3 Super, SEA-LION v4.8 120B |
| 512 GB Mac Studio | M3 / M5 Ultra | 512 GB | 400-700B MoE | Mistral Large 3, Nemotron 3 Ultra, GLM-5.3-Flash, MiniMax M3 |
| One 8-GPU node | 8x H100 / H200 / B200 | 640 GB-1.4 TB | 1T-class MoE at FP8 or INT4 | Kimi K2.6, DeepSeek V4 Pro, GLM-5.3, MiMo-V2.6-Pro, Inkling |
| Cluster | Multi-node, or B300-class | 1.4 TB+ | 2T+ MoE | Kimi K3, Qwen3.8 2.4T |
The two middle tiers are single-user machines. They run a 120B or 700B model for one developer or a small team, not for hundreds of concurrent users. Production concurrency still means data-centre GPUs.
The mixture-of-experts trap still applies at every tier. "17B active" or "41B active" describes compute per token, not memory. Every parameter has to sit in memory so the router can reach any expert. Size the hardware on total parameters.

The model families, one by one
The order reflects fit for enterprise deployment (licence clarity, provenance, track record and support), not benchmark rank. On raw benchmarks, the Chinese families lead. Benchmark numbers below are vendor-reported unless marked otherwise, and every lab uses its own harness. Treat them as direction, not ranking. Hosting options list providers confirmed in the vendor documentation reviewed for this article; all of these models can also be self-hosted.
Mistral (France)
Paris-based, founded in 2023 by former Google DeepMind and Meta researchers. Europe's leading model lab and the main open-weight option for buyers who need EU origin.
- Models: Mistral Large 3 (675B, 41B active), Mistral Small 4 (119B, 6.5B active), Ministral 3 (3B to 14B), Devstral Small 2 (24B coding), Mistral Medium 3.5 (128B dense), plus Voxtral (speech), Shieldstral (guard model) and Leanstral (formal proofs).
- Strengths: the broadest European open line, covering text, code, speech, safety and formal verification. Most of it is Apache 2.0. Mistral Large 3 has the lowest hosted frontier-tier price in this article: $0.50 / $1.50 per million tokens. The strongest enterprise track record of any open-weight lab, with on-premises deployments at BNP Paribas, CMA CGM, Stellantis and the French armed forces.
- Weaknesses: on coding benchmarks, Mistral's open models sit a tier below the top Chinese ones. Its strongest coders, Medium 3.5 and Devstral 2, ship under a modified MIT licence that grants no rights if "the global consolidated monthly revenue of your company (or that of your employer) exceeds $20 million". Larger companies need a commercial licence from Mistral or use the hosted API. Some products have no openly licensed weights: Mistral OCR is API-only, and Codestral weights carry the Mistral Non-Production License, so commercial self-hosting needs an agreement with Mistral. Training data is undisclosed.
- Use cases: multilingual enterprise assistants, RAG, document work, speech, on-device, and any workload that must stay inside EU jurisdiction.
- Where to run it: Mistral La Plateforme (EU), AWS Bedrock, Azure AI Foundry (including the EU Data Zone for Large 3 and Medium 3.5), Google Vertex, IONOS, T-Systems.
Nemotron (NVIDIA, US)
NVIDIA's own models, built to show off and sell its inference stack.
- Models: Nemotron 3 Ultra (550B, 55B active), Nemotron 3 Super (120B, 12B active), Nemotron 3 Nano and 3.5 Lightning (30B, 3B active).
- Strengths: the most transparent large model. NVIDIA publishes the pre- and post-training data. Ultra uses the permissive OpenMDW licence. Hybrid Mamba architecture with 1M context. NVIDIA signs every model on NGC. Ultra is the strongest US open model for reasoning and RAG. Singapore's latest SEA-LION and Malaysia's ILMU-Nemo build on Nemotron.
- Weaknesses: Super and Nano use the NVIDIA Open Model Licence, which terminates if you bypass its guardrails without a replacement. Part of the post-training data was distilled from DeepSeek, Qwen and gpt-oss outputs, per NVIDIA's own card. The small models trail Qwen on coding.
- Use cases: RAG, reasoning, sovereign base models, NVIDIA-standardised infrastructure.
- Where to run it: AWS Bedrock (Super), NVIDIA NIM containers on your own GPUs.
gpt-oss (OpenAI, US)
OpenAI's first open-weight language-model release since GPT-2, published August 2025.
- Models: gpt-oss-120b (117B, 5.1B active), gpt-oss-20b.
- Strengths: the most widely hosted open model. Apache 2.0. The 120b fits a single 80 GB GPU or a 128 GB desktop, and is very cheap hosted ($0.15 / $0.60). Strong tool use. A popular fine-tune base (Japan's GPT-OSS-Swallow).
- Weaknesses: no refresh in 2026. 128K context, text only.
- Use cases: tool-calling agents, reasoning per dollar, a fine-tune base, latency-critical work.
- Where to run it: almost anywhere. AWS Bedrock, Azure, Google Vertex, Oracle OCI, Groq (~500 tokens/s), Cerebras (~3,000 tokens/s), Together, Fireworks, OVHcloud, STACKIT.
Gemma (Google, US)
Google DeepMind's open family, small to mid-size.
- Models: Gemma 4 31B, 26B-A4B, 12B, and the on-device E2B and E4B.
- Strengths: moved to Apache 2.0 with version 4. Covers 140+ languages. The best on-device options in this list. The 31B fits one workstation GPU.
- Weaknesses: the capability ceiling of a 31B model. Earlier Gemma versions (1 to 3) run under Google's own terms, which reserve a right to restrict usage, so check which version you deploy.
- Use cases: on-device and edge, multilingual assistants, vision on a single GPU.
- Where to run it: AWS Bedrock, STACKIT, T-Systems.
Command (Cohere, Canada)
Toronto-based and enterprise-only from the start, with a focus on private deployment.
- Models: Command A+ (218B, 25B active).
- Strengths: Cohere's first Apache 2.0 model. Grounded RAG with citations, 48 languages, built for air-gapped deployment. Cohere says it runs on two H100s at 4-bit.
- Weaknesses: 128K context. Earlier Command models are non-commercial (CC-BY-NC), so check the version.
- Use cases: on-premises RAG with citations, multilingual enterprise search.
- Where to run it: Azure AI Foundry (including Japan, Korea, Australia and India), Cohere's own platform.
Granite (IBM, US)
IBM's enterprise-governance-first family.
- Models: Granite 4.2 30B, 8B, 3B; Granite Guardian 4.1.
- Strengths: Apache 2.0. Built for enterprise RAG and tool calling, with long-context retrieval. Granite Guardian checks RAG groundedness and function-call hallucinations. Fits one workstation GPU.
- Weaknesses: not a frontier model (SWE-bench Verified 57).
- Use cases: RAG, extraction, guardrails, regulated environments that value IBM's governance tooling.
- Where to run it: IBM watsonx.ai, self-hosted.
Inkling (Thinking Machines, US)
The lab founded by former OpenAI CTO Mira Murati. Inkling, released July 2026, is its first large open model.
- Models: Inkling (975B, 41B active, text, image and audio).
- Strengths: the largest Apache 2.0 model from a US lab. Positioned as a base for customisation, with Thinking Machines' Tinker fine-tuning service.
- Weaknesses: new. Well behind the top Chinese models on long-horizon terminal tasks (Terminal-Bench 2.1: 63.8). Node-class hardware.
- Use cases: a US-origin base for heavy fine-tuning.
- Where to run it: self-hosted; Tinker for fine-tuning.
Llama (Meta, US)
The model that started enterprise open-weight adoption, now in retreat. Meta's flagship Muse Spark (April 2026) is closed. The smaller Muse Glimmer (30B, August 2026) is Apache 2.0.
- Models: Llama 4 Maverick and Scout (April 2025), Llama 3.3 70B.
- Strengths: the longest production track record and the deepest tooling. Llama Guard 4 remains a standard guard model.
- Weaknesses: outclassed on capability. The community licence caps use at 700M monthly users and withholds multimodal rights from EU-domiciled companies.
- Use cases: existing deployments, guardrails, conservative environments.
- Where to run it: AWS Bedrock, Azure, Google Vertex, Oracle OCI, Groq, OVHcloud, STACKIT, IONOS.
SEA-LION (AI Singapore)
Singapore's national model programme, part of a S$70M multimodal LLM effort.
- Models: Nemotron-SEA-LION v4.8 (120B, 12B active, on an NVIDIA Nemotron base), Qwen-SEA-LION v4.5 27B.
- Strengths: the best open option for Southeast Asian languages, including Burmese, Filipino, Malay, Tamil, Thai and Vietnamese. MIT licence. 262K context.
- Weaknesses: built on foreign bases (Nemotron, Qwen), so it inherits their licence and provenance questions. Smaller tooling community.
- Use cases: customer-facing work in Southeast Asian languages.
- Where to run it: self-hosted.
Apertus (Swiss AI) and OLMo (AI2, US)
The fully open models: weights, training data, code and recipes all published.
- Models: Apertus 1.5 70B (EPFL, ETH Zurich, CSCS), Olmo 3.1 32B (Allen Institute for AI).
- Strengths: maximum transparency, which makes audits, provenance documentation and EU AI Act records straightforward. Apache 2.0.
- Weaknesses: capability below the frontier open models.
- Use cases: audit-heavy environments, research baselines, public sector.
- Where to run it: self-hosted.
Qwen (Alibaba, China)
Alibaba's model family, and the most adapted open model in the world: Hugging Face counted 151,448 Qwen derivatives by August 2026, and Singapore's SEA-LION, Korea's SK Telecom A.X 4.0 and Thailand's Typhoon all build on Qwen bases.
- Models: Qwen3.8-27B (dense, single GPU), Qwen3.6-35B-A3B (fast local MoE), Qwen3.8-Flash-Next (125B MoE), Qwen3.8 2.4T (open "Max" class).
- Strengths: the widest size range of any family. Qwen3.8-27B is the strongest model that fits one workstation GPU, covering coding, documents and vision. Qwen3Guard is a capable multilingual guard model.
- Weaknesses: PRC origin, which some regulated buyers rule out. Licences vary by checkpoint: the 27B is Apache 2.0, while Flash-Next and the 2.4T need a separate licence if you resell model access at scale. Alibaba also keeps some flagships API-only (Qwen3.7 Max and Plus).
- Use cases: local coding agents, document and vision extraction, a base for regional fine-tunes.
- Where to run it: Alibaba Model Studio (Singapore, Tokyo, Hong Kong, Frankfurt), AWS Bedrock, Together, OVHcloud, STACKIT, IONOS, Cerebras (Qwen3.8-27B).
DeepSeek (China)
A Hangzhou lab funded by the quant fund High-Flyer, known for releasing frontier-class models under MIT with no strings.
- Models: DeepSeek V4.1 Flash (552B total, 16B active, 1M context), V4 Pro (1.6T).
- Strengths: the top open agentic coder on vendor numbers (Terminal-Bench 2.1: 90.6). MIT licence. Low hosted prices: V4 Pro at $1.30 / $2.60 per million tokens on DeepInfra. NIST CAISI rated V4 Pro the most capable PRC model it has tested.
- Weaknesses: the most scrutinised family. CAISI found DeepSeek R1-0528 agents far easier to hijack than US models, and censorship persists when R1 runs locally. Several governments restrict DeepSeek's apps, and some ban wording covers "products". Large hardware footprint: node-class at minimum.
- Use cases: coding agents and long-context work where origin is acceptable and guardrails are in place.
- Where to run it: Azure AI Foundry (V4), AWS Bedrock (V3.2), Alibaba Model Studio (V4), Together, DeepInfra.
Kimi (Moonshot AI, China)
A Beijing lab focused on long context and agentic workloads.
- Models: Kimi K2.6 and K2.7 Code (1T total, 32B active), Kimi K3 (2.8T, 104B active).
- Strengths: K3 leads open models on reasoning (GPQA 93.5) and document understanding. K2.6 runs on one 8x H100 node at native INT4 and is strong at agentic coding and multi-agent work.
- Weaknesses: K3 needs a cluster. K3 ships under a custom licence that requires a separate agreement for model-as-a-service operators above $20M revenue over any consecutive 12 months. CAISI found Kimi K2 Thinking heavily censored in Chinese, much less in English.
- Use cases: agent swarms, coding, long-document reasoning.
- Where to run it: Azure AI Foundry (K2.5 to K2.7), AWS Bedrock, Alibaba Model Studio (K3), Together, DeepInfra.
GLM (Z.ai, China)
Z.ai, formerly Zhipu, a Tsinghua University spin-out.
- Models: GLM-5.3 (744B, 40B active), GLM-5.3-Flash (320B, 18B active).
- Strengths: strong at coding and cyber tasks. GLM-5.3-Flash delivers near-flagship coding at a fraction of the price, under MIT, and fits a 512 GB Mac Studio.
- Weaknesses: GLM-5.3 added a licence clause requiring a Z.ai security review for model-as-a-service firms above $10B aggregate revenue over any consecutive 12 months. GLM-5 had no such clause, which shows why licence review has to happen per version.
- Use cases: coding agents, security tooling, cost-sensitive reasoning.
- Where to run it: AWS Bedrock (GLM 5), Alibaba Model Studio, Together, Fireworks, Z.ai's own API.
MiMo (Xiaomi, China)
Xiaomi's AI lab, new to frontier models but moving fast.
- Models: MiMo-V2.6-Pro (1T, 42B active, text, image, video and audio).
- Strengths: the highest open-model score on the independent Artificial Analysis index (46). MIT licence. Published RL training environments.
- Weaknesses: little enterprise track record yet. Node-class hardware.
- Use cases: general assistant and multimodal agents where capability matters most.
- Where to run it: self-hosted; check current hosted availability.
MiniMax (China)
A Shanghai lab focused on long-context agents.
- Models: MiniMax M3 (428B, 23B active, 1M context).
- Strengths: efficient sparse attention for very long agent sessions.
- Weaknesses: the most restrictive licence in this list. It requires "Built with MiniMax M3" attribution, written authorisation above $20M a year, and bans military use. The earlier M2.7 is non-commercial.
- Use cases: long-context agent work, after legal review.
- Where to run it: Together (M3); AWS Bedrock and Google Vertex (earlier M2 versions).
The regional long tail
- Korea: LG K-EXAONE 2.0 (750B, Apache 2.0, Korean plus 9 languages), Upstage Solar Open2 (250B, Korean, Japanese, English), Naver HyperCLOVA X SEED, Kakao Kanana-2. Korea's national programme requires from-scratch training.
- Japan: LLM-jp-4.1 (33B, Apache 2.0, from the National Institute of Informatics), GPT-OSS-Swallow-120B (a Japanese fine-tune of gpt-oss), Rakuten AI 3.0 (on a DeepSeek-V3 architecture).
- India: Sarvam-105B and 30B (Apache 2.0, 22 Indian languages, trained from scratch).
- Thailand: Typhoon 2.5 (Apache 2.0).
- Other Chinese labs: Tencent Hy4-preview, Baidu ERNIE 4.5, ByteDance Seed-OSS-36B, Ant Group Ling-3.0, StepFun Step-3.7-Flash.
- US niche: Arcee Trinity, Microsoft Phi-4 (small multimodal), Liquid AI LFM2 (on-device; licence capped above $10M revenue), ServiceNow Apriel.
- UAE: TII Falcon-H1R (small hybrid models).
Best pick per use case
| Use case | On a workstation (up to 128 GB) | On one enterprise node | If origin must be outside China |
|---|---|---|---|
| Coding agent | Qwen3.8-27B, Devstral Small 2 | DeepSeek V4.1 Flash, GLM-5.3, Mistral Medium 3.5 (licence cap) | Devstral Small 2, Mistral Small 4, Inkling |
| General assistant | Gemma 4 31B, Mistral Small 4 | Kimi K2.6, MiMo-V2.6-Pro, Mistral Large 3 | Mistral Large 3, Command A+, Nemotron 3 Ultra |
| RAG and extraction | Granite 4.2 30B, Qwen3.8-27B, Mistral Small 4 | Command A+, Mistral Large 3 | Command A+, Mistral Large 3, Granite, Nemotron 3 Ultra |
| Southeast Asian languages | Qwen-SEA-LION v4.5 27B | Nemotron-SEA-LION v4.8 | SEA-LION v4.8, Gemma 4 |
| Documents and vision | Qwen3.8-27B, Gemma 4, Mistral Small 4 | Kimi K3, Mistral Large 3 | Gemma 4, Mistral Small 4, Mistral Large 3 |
| Speech | Voxtral Mini 3B, Voxtral Realtime 4B | Voxtral Small 24B | Voxtral family |
| Guardrails | Shieldstral 1.0, Qwen3Guard, Granite Guardian 4.1 | Llama Guard 4 | Shieldstral, Granite Guardian, Llama Guard |
| Formal verification | Leanstral 1.5 | Leanstral 1.5 | Leanstral |
| On-device | Gemma 4 E2B / E4B, Ministral 3 | n/a | Gemma 4, Ministral 3, Phi-4 |
| Maximum transparency | Apertus 1.5, Olmo 3.1 | Nemotron 3 Ultra | All three |
Among European labs, Mistral is the only one with an open model in most rows, including speech and formal proofs. Its Apache 2.0 line (Large 3, Small 4, Ministral 3, Devstral Small 2, Voxtral, Shieldstral, Leanstral) is the default shortlist for buyers who need EU origin. One exception inside the family: Voxtral TTS (text-to-speech) is non-commercial.

Where to rent open weights
"Start on a hosted open-weight API" raises the next question: whose? Model catalogues, regions and data defaults differ more than the marketing suggests. Status per vendor documentation, October 3, 2026.
| Vendor | Jurisdiction | Type | Open models (selection) | APAC regions | EU regions | Data defaults |
|---|---|---|---|---|---|---|
| AWS Bedrock | US | Hyperscaler | gpt-oss, DeepSeek V3.2, Qwen3, Kimi K3, GLM 5, MiniMax M2.5, Mistral Large 3, Llama 4, Gemma 4, Nemotron 3 Super | In-region in Jakarta, Tokyo, Mumbai, Sydney, Melbourne | Frankfurt, Stockholm, Milan, Ireland, London (per model) | Zero retention configurable per region, enforceable org-wide. Model vendors see no prompts |
| Azure AI Foundry | US | Hyperscaler | DeepSeek V4, Kimi K2.5-K2.7, Llama 4, Mistral Large 3, gpt-oss-120b | Australia East, Japan East/West, Korea Central, South India; global deployment, no APAC data zone | EU Data Zone for DeepSeek V4-Flash, Mistral Large 3, Mistral Medium 3.5 | Global deployments may process in any Azure region |
| Google Vertex AI | US | Hyperscaler | DeepSeek V3.2, Qwen3, gpt-oss, Llama 4, Kimi K2, MiniMax M2, GLM 5.2, Mistral | Processing: gpt-oss and Qwen3 in the US, DeepSeek V3.2 global. No Asian region | Same: US or global processing | DeepSeek V3.2 and Qwen3-235B endpoints retire October 21, 2026 |
| Oracle OCI | US | Hyperscaler | Llama 4, Llama 3.3, gpt-oss | Osaka on-demand, Hyderabad dedicated | Frankfurt | Dedicated clusters available |
| Alibaba Cloud Model Studio | China (parent) | APAC cloud | Qwen 3.x, DeepSeek V4, Kimi K3, GLM 5.3 | Singapore, Tokyo, Hong Kong | Frankfurt | No training on customer data. Inference can route globally unless you pick a geo-bounded scope |
| Fireworks AI | US | Inference specialist | Broad open catalogue | Dedicated: Malaysia, Tokyo, New South Wales | Dedicated: Frankfurt, Iceland | Zero retention by default for open models |
| Together AI | US | Inference specialist | DeepSeek V4.1, Kimi K3, GLM-5.3, MiniMax M3, Qwen 3.x, gpt-oss | No region choice on serverless | Via dedicated endpoints or VPC | Stores prompts by default; zero retention is opt-in |
| DeepInfra | US | Inference specialist | Broad open catalogue | Not stated | Not stated | In-memory only, no training |
| Groq | US | Speed specialist (~500 tokens/s on gpt-oss-120b) | gpt-oss, Llama 3.x, previews | None | None | No retention by default; stored data sits in US buckets |
| Cerebras | US | Speed specialist (~3,000 tokens/s on gpt-oss-120b) | gpt-oss-120b, Qwen3.8-27B | Not stated | Not stated | SOC 2 Type 2 |
| Baseten | US | Inference platform | Any model, custom weights | Region restrictions configurable | Region restrictions configurable | Zero retention by default; can run inside your own VPC. SOC 2 Type II, HIPAA |
| OpenRouter | US | Aggregator | Everything, routed to other hosts | Depends on upstream host | EU in-region routing on business plans | Per upstream host; can filter out hosts that train on data |
| Hugging Face Inference Providers | US / France | Router | Routes to Fireworks, Groq, Together, OVHcloud, Scaleway and others | Depends on provider | Pin EU providers | Provider picked dynamically unless pinned |
| Mistral La Plateforme | France | EU sovereign | Mistral Large 3, Small 4, Ministral 3 | n/a | EU | Enterprise accounts opted out of training; zero retention available |
| OVHcloud AI Endpoints | France | EU sovereign | Llama 3.3, Qwen 3.x, gpt-oss | n/a | EU | Not stated on catalogue |
| IONOS AI Model Hub | Germany | EU sovereign | Llama, Mistral, Qwen 3.5 | n/a | Germany | Not stated on catalogue |
| STACKIT AI Model Serving | Germany (Schwarz Group) | EU sovereign | gpt-oss, Llama 3.3, Qwen3-VL, Qwen3.8-27B, Gemma | n/a | EU01 | Not stated on catalogue |
| T-Systems AI Foundation Services | Germany | EU sovereign | Mistral, Qwen, Gemma | n/a | T-Cloud Germany; some models served via Azure Sweden or Google Cloud | ISO 27001 |
Also in the market, not covered here: Scaleway, Nebius, Aleph Alpha, BytePlus ModelArk, Huawei Cloud, Tencent Cloud, Baidu Qianfan, and local GPU clouds such as Singtel RE:AI, Sakura Internet, Naver Cloud and Yotta, which mostly sell GPU capacity rather than managed model APIs.
Six things the table shows.
Check where inference actually runs, not where the endpoint lives. Region lists tell you where you can create an endpoint. They do not always tell you where your data is processed. AWS Bedrock runs its open models in-region, in Jakarta, Tokyo, Mumbai and Sydney among others. Azure offers them as global deployments that may process in any Azure region. Google processes gpt-oss and Qwen3 in the US and DeepSeek V3.2 globally. Alibaba Cloud serves Qwen, DeepSeek, Kimi and GLM from Singapore, Tokyo, Hong Kong and Frankfurt, and needs a geo-bounded scope to keep inference in-region. If a customer needs processing in a specific country, choose the host by its processing location, or deploy the weights yourself on GPU instances in that region.
Model origin and data jurisdiction are separate choices. You can run DeepSeek, Qwen or Kimi on AWS in Frankfurt, or Qwen and gpt-oss on STACKIT in Germany. AWS documents that model providers have no access to prompts or completions. With open weights, the origin risk stays in the weights and the licence. It leaves the data path.
Every host carries some jurisdiction. US-headquartered hosts fall under the US CLOUD Act, including the specialists and routers. Alibaba swaps that for PRC-law exposure through its parent. The EU-headquartered options (Mistral, OVHcloud, IONOS, STACKIT, T-Systems for its German-hosted models) sit outside both. For French sensitive state data under SecNumCloud, that difference decides eligibility.
Retention defaults differ. Fireworks, DeepInfra, Baseten and Groq keep nothing by default. Together stores prompts unless you opt out. Read the data-handling page before the pricing page.
Hosted catalogues churn too. Bedrock commits to keeping a model at least 12 months after launch. Google deprecated its DeepSeek V3.2 and Qwen3-235B endpoints in July 2026, with retirement on October 21, three months later. Specialist serverless catalogues turn over faster still, and preview tiers are evaluation-only. The pinning argument applies to hosted open models as well: if a workload must not change, use a dedicated deployment or your own weights. The difference from a closed API is that the exit stays open when the host drops the model.
Hyperscaler or specialist is a trade. Hyperscalers bring existing contracts, IAM, private networking and committed spend, but lag on new models: Bedrock serves DeepSeek V3.2 while Azure and Alibaba already offer V4. Specialists ship new models within days, are cheaper, and offer more dedicated regions, but publish thinner compliance evidence. Bring-your-own fine-tunes are geographically narrow on hyperscalers (AWS Custom Model Import runs in Frankfurt and three US regions; Azure open-model fine-tuning is US-only).
Sources
Models and hosting
- Qwen3.6-27B release post (April 22, 2026)
- HackerNews discussion of the Qwen3.6-27B release (April 22, 2026)
- Alibaba Cloud Model Studio, international (Singapore) endpoint (accessed April 23, 2026)
- Pellet OWL 2 reasoner on GitHub
- Artificial Analysis open-weights leaderboard (accessed October 3, 2026)
- DeepSWE leaderboard (September 22, 2026)
- Qwen3.8-27B model card (August 2026)
- DeepSeek V4.1 Flash release note (September 10, 2026)
- Kimi K3 model card and licence (July 2026)
- GLM-5.3 model card and licence (August 2026)
- MiniMax M3 model card and licence (June 2026)
- Mistral Large 3 (December 2025)
- Mistral Small 4 (March 2026)
- Mistral Medium 3.5 model card and licence (April 2026)
- gpt-oss-120b (August 2025)
- Google: Gemma 4 (April 2, 2026)
- NVIDIA Nemotron 3 Ultra (June 2026)
- Linux Foundation: OpenMDW-1.1 (May 28, 2026)
- Thinking Machines Inkling (July 2026)
- Cohere Command A+ (May 2026)
- IBM Granite 4.2 30B (August 2026)
- Apertus 1.5, ETH AI Center (July 2026)
- Nemotron-SEA-LION v4.8 announcement (September 2026)
- Llama 4 acceptable use policy
Economics and adoption
- Menlo Ventures: 2025 State of Generative AI in the Enterprise (December 9, 2025)
- OpenRouter and a16z: State of AI (December 2025)
- Hugging Face: State of Open Models, Summer 2026 (August 14, 2026)
- UC Today on Mozilla's State of Open Source AI 2026 (July 20, 2026)
- Anthropic pricing, OpenAI pricing, Mistral pricing, Together pricing, Groq models, Alibaba Model Studio pricing, Lambda pricing (all accessed October 3, 2026)
- AIMultiple GPU price index (September 2026)
- vLLM: gpt-oss support (August 5, 2025)
- Spheron: Mistral API pricing vs self-hosted (July 23, 2026)
- CMA CGM adopts Mistral AI (April 2025)
- FinTech Futures: BNP Paribas and Mistral AI (July 2024)
- TechRepublic: Mistral and the French armed forces (January 2026)
- Stellantis and Mistral AI expand collaboration (October 2025)
- Orange and Mistral AI (February 2025)
- Deutsche Telekom: Germany's first AI factory for industry (February 2026)
- GoTo and Indosat: Sahabat-AI 70B (June 2, 2025)
- SK Telecom A.X 4.0 (July 2025)
- Rakuten AI 3.0 (March 17, 2026)
- Fintech News: OCBC generative AI (2023)
- Next: Albert not generalised (January 14, 2026)
- MDDI: S$1bn National AI R&D Plan (January 24, 2026)
- Korea Herald: national AI model first-round results (January 15, 2026)
- Sarvam 30B and 105B (February 2026)
Hosting providers (all accessed October 3, 2026)
- AWS Bedrock: models at a glance, gpt-oss-120b model card, data protection, Custom Model Import, regional availability by model
- Azure AI Foundry: region availability, deployment types
- Google Vertex AI open models, gpt-oss-120b, DeepSeek-V3.2, Qwen3 235B
- Oracle OCI Generative AI models by region
- Alibaba Model Studio regions, privacy notice
- Fireworks regions, data handling
- Together privacy and security
- DeepInfra data privacy
- Groq data handling
- Cerebras models, trust center
- Baseten security
- OpenRouter privacy and logging
- Hugging Face Inference Providers
- Mistral API data storage
- OVHcloud AI Endpoints catalogue
- IONOS AI Model Hub models
- STACKIT shared models
- T-Systems LLM Hub
Security and control
- JFrog: malicious models on Hugging Face (March 2024)
- ReversingLabs: nullifAI (February 2025)
- CVE-2025-32434 advisory (April 2025)
- HiddenLayer: malware in trending Hugging Face repository (May 7, 2026)
- OpenSSF Model Signing v1.0 (April 4, 2025)
- NIST CAISI evaluation of DeepSeek models (September 30, 2025)
- NIST CAISI evaluation of Kimi K2 Thinking (December 12, 2025)
- NIST CAISI evaluation of DeepSeek V4 Pro (May 1, 2026)
- R1dacted: censorship in DeepSeek R1 (May 2025)
- The Hacker News on CrowdStrike's DeepSeek-R1 code findings (November 2025)
- Anthropic: a small number of samples can poison LLMs of any size (October 9, 2025)
- Euronews: Czech Republic bans DeepSeek in state administration (July 10, 2025)
- NVIDIA Open Model License (October 24, 2025)
- OpenAI API deprecations
- Anthropic model deprecations
- Gemini API deprecations
- Simon Willison on Anthropic's postmortem (September 17, 2025)
- Confidential computing overhead on NVIDIA B200 (August 2026)
Singapore
- MAS Consultation Paper on Guidelines on AI Risk Management (November 13, 2025)
- MAS written reply on agentic AI in financial services (August 5, 2026)
- IMDA Model AI Governance Framework for Agentic AI, v1.5 (May 20, 2026)
- IMDA Starter Kit for Testing LLM-Based Applications (January 2026)
- GovTech GenAI control (accessed October 3, 2026)
- MAS Notice 658 on Management of Outsourced Relevant Services
EU
- Commission guidelines for providers of general-purpose AI models (July 2025)
- White & Case: EU AI Omnibus enters into force (July 2026)
- ESAs designate critical ICT third-party providers under DORA (November 18, 2025)
- EU Data Act explained
APAC residency
- OJK Regulation 11/2022 on IT for commercial banks
- BNM Outsourcing policy document (October 2019)
- BOT Notification FPG 19/2559
- Vietnam PDP Law 91/2025/QH15
- RBI Master Direction on Outsourcing of IT Services (April 2023)
- APRA Prudential Standard CPS 230
- My Health Records Act 2012, s77
- Kim & Chang on Korean financial cloud rules
- Taylor Wessing on §393 SGB V
