Open weights don't make enterprise AI cheaper. They make it portable.

Open weights don't make enterprise AI cheaper. They make it portable.

Take a hypothetical case. You run a software startup. Your product puts an AI agent to work on customers' business data: pricing models, product catalogues, approval rules. In development it runs on a closed frontier API, and it works.

Then the pipeline fills with the customers you actually want: a bank in Jakarta, an insurer in Singapore, a healthcare group in Germany, a manufacturer in Australia. Each sends a vendor due-diligence questionnaire, and the same four questions show up in every one:

  • Which model processes our data, and in which country?
  • Will you tell us before that model changes?
  • Can our regulator audit your providers?
  • What happens if your model provider withdraws the service?

Then each market adds its own twist. The Indonesian bank's regulator expects its systems processed in Indonesia unless approved otherwise. The German healthcare group needs processing in the EU on certified infrastructure. The Australian customer has to notify its regulator before data goes offshore. None of this is the customer being difficult. It is their regulators, and it flows down to you as their vendor.

With one closed API, the honest answers are thin: the provider's current model, wherever the provider runs it; only when the provider tells us; probably not; we move to whatever the provider offers next. And you cannot place that model in Jakarta and Frankfurt at the same time.

An open-weight model changes the answers. You can name the exact version, pin it, run the same weights in each region your customers need, and move them if a host fails you. One product, one set of evals, many jurisdictions. That is what this article is about.

It is not about the leaderboard. Benchmark leads now expire in months: Qwen's 27B beat its own 397B flagship in April and was overtaken by its successor in August. It is not mainly about price either. Closed providers cut prices too, and a choice made only on price gets reversed at the next price cut.

The position: open weights turn the model from a vendor relationship into a component you control. They don't make enterprise AI cheaper. They make it portable. And portability is what regulators in Singapore and the EU are now writing into their rules.

The short version:

  • Cost: open vs closed is a cost decision. Hosted open models are 13 to 17 times cheaper per token at the frontier and mid tiers. Hosted vs self-hosted is a control decision; owning GPUs rarely beats renting the same open model.
  • Default path: start on a hosted open-weight API. Move a workload in-house when one of three triggers fires: a rule that requires local processing, steady high volume, or fine-tuning.
  • Regulators: MAS, IMDA, the EU AI Act and DORA do not prefer open or closed. They ask for models you can test, audit, pin and exit. Open weights make those answers easy.
  • This quarter: separate the model choice from the hosting choice, classify workloads by residency rule, build your own eval set, set up a provenance check before the first download, and test the exit swap once a quarter.

In this article:

Part 1: The argument

Part 2: Reference


Part 1: The argument

What open weights actually change

Four things shift when you hold the weights, not just an API key.

Closed API Open weights
Exit Model retired or repriced on the provider's schedule Same weights run on another host, or your own hardware, unchanged
Data residency Wherever the provider's region map allows Wherever you put the GPU
Cost behaviour Per token, scales linearly with usage Per GPU-hour when self-hosted, flat once the hardware is busy
Customisation Whatever fine-tuning the provider exposes LoRA or full fine-tune on your own data, and you keep the result

Exit is the one most teams underrate. The same open-weight model is served today by several independent inference providers. If one raises prices, degrades latency or drops the model, you move the endpoint. Your prompts, evals and agent code stay the same. No closed model gives you that.

Residency becomes an infrastructure decision, not a contract negotiation. You do not ask a provider whether they can process data in Jakarta. You deploy in Jakarta.

Cost becomes predictable at scale. API spend tracks usage one-to-one, so a successful rollout is also a budget surprise. A self-hosted node costs the same at 30% utilisation as at 90%. That is worse for spiky, low-volume work and much better for steady, high-volume work.

Customisation becomes an asset you own. A LoRA adapter trained on your contracts, tickets or product catalogue is a file on your storage. It does not disappear when a vendor deprecates the base model.

So when you evaluate an open-weight model, ask about portability first. Benchmark rank comes second.

Latest model or predictable model

The instinct is to run the latest and greatest. For some work that is right: a coding assistant or a research agent gets better every time the model does, and nobody minds if it answers differently next week.

Many enterprise use cases want the opposite. An agent that extracts fields from contracts, prices a quote or changes a customer's system configuration has to behave the same way tomorrow as it did on the day it was validated. Every model change means re-running the evals, re-checking the edge cases and, in a regulated customer, telling the regulator's control framework why the output changed. A few benchmark points are not worth that.

Latest model Pinned model
Typical workloads Coding assistants, research, drafting, exploration Extraction, classification, pricing, configuration changes, regulated decisions
What matters Capability ceiling Same input, same output; validated behaviour
Model change Welcome A release with regression testing and sign-off
Who sets the upgrade date Whoever ships the newest model You, after your evals pass

Closed APIs serve the left column well. The right column is where they struggle: the provider retires versions on its schedule, and behaviour can shift under an unchanged model name. Open weights let you run both columns at once. Use the newest model where it pays, and keep a pinned, validated version for the workloads that have to stay predictable, upgrading them when you decide.

Deployment control is the point, not the licence

Look at what the largest Mistral enterprise deals actually buy. BNP Paribas, CMA CGM and the French armed forces run Mistral's commercial models on their own or French infrastructure, not necessarily the Apache-licensed weights. What they pay for is the right to run the model where they decide, under their own controls. The open licence is one way to get that right. A commercial on-premises licence is another. Either way, the sovereignty comes from deployment control, and that is what your regulated customers are asking about.

Your evals are the asset you carry

Models change every quarter. Hosts come and go. What survives every change is your evaluation set: the few hundred real tickets, contracts, quotes or configuration tasks with known correct outcomes, scored the same way every time.

That set is what makes the rest of this article work. It tells you whether an open model is good enough to replace a closed one. It tells you whether a new version can replace the pinned one. It gives a regulator evidence that behaviour did not change after a host switch. Building a demo on a frontier model is easy. Proving a production workload still behaves after the model underneath changes is the hard part, and only your own evals do that.

Treat the eval set as a product asset with an owner, version control and a release gate. It is the one piece of the AI stack you fully own, whichever model you run.

The interface layer makes the swap cheap

Portability only pays if switching models is cheap. Two things make it cheap today. Most hosts, open and closed, expose an OpenAI-compatible API, and many also accept the Anthropic API format. And agent tooling increasingly connects through the Model Context Protocol (MCP), so tools and data sources are wired to the agent, not to a specific model.

Build on those two layers and changing the model under an agent is a configuration change plus an eval run. Build on one provider's proprietary features (its own agent runtime, assistants API or fine-tuning format) and every switch becomes a migration project. The portability argument is decided in your architecture long before it reaches a procurement questionnaire.

What it costs: open is cheaper, owning the GPU often is not

The title says open weights do not make AI cheaper. That needs precision, because two different comparisons get mixed up.

Open model vs closed model: open wins on price. List prices per million tokens, fetched October 3, 2026:

Tier Open-weight, hosted (input / output) Closed (input / output) Output price gap
Frontier Mistral Large 3: $0.50 / $1.50. DeepSeek V4 Pro: $1.30 / $2.60. GLM-5.3: $1.40 / $4.40 Claude Opus 5.5: $4 / $20. GPT-6-astra: $10 / $50 ~13x (Large 3 vs Opus 5.5)
Mid gpt-oss-120b: $0.15 / $0.60. DeepSeek V4.1 Flash: $0.30 / $1.20 Claude Sonnet 5.5: $2 / $10. GPT-6.1-sol: $2 / $10 ~17x (gpt-oss-120b vs Sonnet 5.5)
Small gpt-oss-20b: $0.075 / $0.30. Qwen3-30B-A3B (Alibaba, Singapore region): $0.20 / $0.80 Claude Haiku 4.5: $1 / $5. GPT-6-luna: $0.10 / $0.50 ~17x vs Haiku 4.5; GPT-6-luna is priced like an open model

At the small end the gap is closing: OpenAI prices GPT-6-luna at $0.10 / $0.50, in open-model territory. Same-tier comparisons also flatter open models, because closed models still lead on quality. Mozilla's 2026 open-source AI report puts the gap at roughly 6x at comparable quality. Still large.

Self-hosting vs renting the same open model: renting usually wins. Hosted open-weight providers run at high utilisation and batch across customers. You will not beat that on a lightly used GPU.

A worked estimate, with every assumption visible:

Assumption Value
Model gpt-oss-120b on one H100, vLLM, ~2,000 output tokens/s aggregate (±50%)
Rented H100 $3.99/hour list price, ~$2,900/month
Owned 8x H100 in a Singapore colo ~$15,600/month: server amortised over 36 months, power, rack, half an engineer
Input:output ratio 3:1
Scenario Cost per million output tokens
Rented H100 at 100% utilisation $0.55
Rented H100 at 50% $1.11
Rented H100 at 25% $2.22
Hosted gpt-oss-120b API $0.60
Claude Sonnet 5.5 API $10.00

A rented GPU only matches the hosted open API at near-full utilisation. The owned 8-GPU node breaks even against the hosted open API at roughly 35% average utilisation, and at about 70% if real throughput is half the assumption. Published analyses land in the same place: a Spheron study of Mistral Large 3 on 8x H200 concluded the cluster could not reach break-even volume against the API, and justified self-hosting on residency, not cost.

So the economics split cleanly:

  • Open vs closed is a cost decision. If an open model is good enough for the task, the hosted version is an order of magnitude cheaper.
  • Hosted vs self-hosted is a control decision. You pay for it unless your volume is steady enough to keep GPUs busy, or a residency rule leaves no choice.

That is the portability argument in numbers. Start hosted, and move only when control is worth the premium.

What regulators now ask for

None of the new rules prefers open or closed models. What they prefer is a model you can test, audit, pin and leave. Read them closely and that is a list of open-weight properties.

Instrument Status (Oct 2026) What it asks of a firm using third-party AI
MAS Guidelines on AI Risk Management (Singapore, all financial institutions) Consultation Nov 2025; MAS said in Aug 2026 the final text is coming "soon", with a proposed 12-month transition AI inventory including third-party models. Test the vendor model on your own use case and data. If the vendor discloses too little, limit use under §4.3 or test compensatorily under §4.11. For open-source models, check provenance and training-data integrity (§4.11(c)). Assess concentration on key providers, plan for a vendor dropping support, secure audit rights and change notifications (§4.11(d)-(f))
IMDA Model AI Governance Framework for Agentic AI (Singapore) v1.0 Jan 2026, v1.5 May 2026; voluntary Limited visibility and control over third-party components is a named risk factor. Require disclosures, scoped credentials and tool-call logging, or reassess the deployment
IMDA Starter Kit for Testing LLM Applications (Singapore) Jan 2026; voluntary Pre-deployment testing for hallucination, bias, harmful content, data leakage and adversarial prompts. Warns explicitly about backdoored fine-tunes from unverified sources
EU AI Act, amended by the Digital Omnibus GPAI duties live since Aug 2025, Commission enforcement since Aug 2026; high-risk obligations moved to Dec 2027 Deployer duties apply whether the model is open or closed. The open-source exemption (Art. 53(2)) only relieves the model's publisher, never your deployment, and never models above 10^25 FLOP
EU Commission GPAI guidelines Jul 2025 You become the "provider" of a modified model only if the modification uses more than a third of the original training compute. A LoRA fine-tune keeps you a deployer
DORA (EU financial sector) Applies since Jan 2025 An LLM API is an ICT third-party service: register entry, contract terms on data location, audit and termination, a tested exit strategy

Three consequences follow.

Opacity has a price, and the deployer pays it. MAS §4.3 and §4.11 are explicit: if your vendor will not tell you enough, you run more tests or you use the model for less. With weights in your own environment, the information gap does not exist.

Silent model updates become a control finding. MAS wants notification of third-party model changes and an impact assessment. Pinned weights do not change unless you change them. With an API, you need version pinning in the contract and regression tests on every release.

Open weights are not a free pass. The same MAS paragraph that makes the exit case also names open-source models with weak security controls as a risk. A Hugging Face download is a software supply-chain dependency. Treat it like one: verified source, checksums, scanning, red-teaming before production.

Residency by industry: where inference has to stay

Most discussions of data residency blur two different rules. "Store locally" means a copy of the data must sit in-country; processing elsewhere can still be allowed. "Process locally" means the compute itself must run in-country. For an LLM, only the second one forces inference into the country.

Market Sector Rule Effect on an LLM
Singapore Government GovTech GenAI control under IM8 Overseas-hosted models: data up to Restricted only. Singapore-hosted: up to Confidential. Provider must not log, store or train on inputs
Singapore Banks MAS Notice 658 on outsourcing (in force Dec 2024) No localisation. Register, due diligence, audit rights for the bank and MAS, exit planning
Indonesia Banks POJK 11/2022 Systems placed in Indonesian data centres; offshore placement needs OJK approval. Process locally by default
Indonesia Insurance, finance, fintech POJK 4/2021 Same pattern as banks
Malaysia Banks, insurers BNM RMiT (reissued Nov 2025), Outsourcing policy Consult BNM before first public-cloud use for critical systems; prior approval for material outsourcing, with processing locations listed
Thailand Banks BOT FPG 19/2559 Public cloud for critical IT needs BOT approval 30 days ahead
Vietnam All data controllers PDP Law 91/2025 (from Jan 2026) Offshore processing allowed, but a transfer impact assessment must be filed within 60 days
India Payments RBI payment data directive (2018) Store locally. Processing abroad allowed, data back in India within 24 hours
India Securities SEBI cloud framework (2023) Storage and processing in empanelled Indian data centres. Process locally
China Banking, payments PBOC rules on personal financial information Stored, processed and analysed in China. Process locally
Korea Financial Electronic Financial Supervisory Regulation, network separation Credit and ID data on cloud must sit in Korea; internet-based LLM services only via sandbox exceptions. Process locally
Australia Banks, insurers, super APRA CPS 230 (Jul 2025) Notify APRA before any material offshoring. Regulator access rights, orderly exit
Australia Health My Health Records Act s77 No holding or processing outside Australia. Process locally
EU Financial DORA Contract must state processing regions; exit strategy; non-EU critical providers need an EU subsidiary
Germany Health §393 SGB V Processing in Germany, EU/EEA or adequate country, plus BSI C5 attestation
France Sensitive state data SREN law + SecNumCloud Providers must be immune to non-EU extraterritorial law. US hyperscaler APIs excluded

Read the table by column, not by row, and three patterns show.

Outright bans are rare. Approval and notification gates are common. Malaysia, Indonesia, Thailand, Australia and the Philippines all let a bank use an offshore service, after the regulator has seen it. The real cost of a closed API in APAC banking is lead time and paperwork per use case.

Regulators keep audit rights over your providers. MAS 658, RBI, APRA CPS 230 and DORA all require that the regulator can inspect the service provider. Frontier-model API terms rarely concede that. A model running inside your own environment needs no such negotiation.

Exit plans are mandatory almost everywhere. RBI, APRA, DORA and MAS all require one. From January 2027 the EU Data Act also removes cloud switching charges. The easiest exit plan to evidence is a model you can redeploy elsewhere tomorrow.

For a group operating across five ASEAN markets, this decides whether one AI platform covers the region or whether each country gets its own exception file. A closed model is available where its provider has a region. An open-weight model is available wherever there is a GPU.

Two more regional factors count. Some regional cloud providers serve open-weight models from Singapore endpoints with OpenAI- or Anthropic-compatible APIs, so you get regional latency without changing agent code. And open-weight families with strong Asian-language training can be fine-tuned further on Bahasa, Thai or Vietnamese corpora you own.

Who is already doing it

The adoption numbers point two ways. Menlo Ventures' December 2025 survey of 495 US enterprises found open-source models at 11% of enterprise LLM use, down from 19% a year earlier. Developer usage moved the other way: open-weight models carried about a third of tokens on OpenRouter by late 2025, and Hugging Face counted 151,448 Qwen derivatives by August 2026. Mozilla's 2026 report names the gap: 51% of developers run open models in production, against 63% for closed.

The gap is operations, not model quality. Where enterprises do commit, sovereignty and customisation are the stated reasons.

Organisation Model Deployment Why
BNP Paribas Mistral commercial models On-premises, all business lines (2024, extended May 2026) EU regulation, data sovereignty
CMA CGM Mistral €100M over five years, Mistral engineers embedded in Marseille Customisation
French Ministry of Armed Forces Mistral Framework agreement 2026-2030, French infrastructure, fine-tuned on defence data Sovereignty
Stellantis Mistral In-car assistant, engineering analytics Customisation
Orange Codestral Hosted on Orange Business infrastructure, resold to business customers Sovereignty
Deutsche Telekom SOOFI, ~100B open European LLM Industrial AI Cloud, Munich, 10,000+ GPUs Sovereignty
GoTo + Indosat (Indonesia) Sahabat-AI, Gemma-based, 70B version 2025 Trained and served in-country on GPU Merdeka; runs in GoPay Local languages, sovereignty
SK Telecom (Korea) A.X 4.0, Qwen2.5 continued pre-training Open weights, local deployment Korean token efficiency
SCB 10X (Thailand) Typhoon, on Mistral, Qwen and Gemma bases Open weights Thai language
Rakuten (Japan) Rakuten AI 3.0, ~700B MoE on a DeepSeek-V3 architecture Apache 2.0 Japanese model, government-subsidised

Two cases keep the picture honest. OCBC built its internal GPT on Azure OpenAI in a controlled environment: a sovereignty-sensitive bank that chose closed. And the French state's Albert assistant, built on Llama and Mistral on government infrastructure, was not rolled out further "in its current form" in January 2026 after a pilot at 48 sites. Open weights do not rescue a weak product.

Sovereign AI in APAC runs on open weights

Country Programme Open weights? Base
Singapore NAIRD, S$1bn+ for 2025-2030; S$70M national LLM programme (SEA-LION, MERaLiON) Yes Foreign open bases: Gemma, Llama, Qwen, now Nemotron
Indonesia Sahabat-AI on GPU Merdeka; Danantara sovereign AI infrastructure push (Sept 2026) Yes Gemma, Llama
Malaysia YTL ILMU ILMU 1.0 closed; ILMU-Nemo-30B on NVIDIA Nemotron Foreign open bases
Thailand ThaiLLM (NSTDA), Typhoon Typhoon yes Qwen, Gemma, Mistral
Japan GENIAC subsidies, ABCI 3.0 Mixed Rakuten on DeepSeek-V3
Korea Independent AI Foundation Model Project Yes (K-EXAONE 2.0, A.X K2 under Apache 2.0) From scratch by rule: Naver was cut in January 2026 for using foreign weights
India IndiaAI Mission, ~US$1.1bn Yes (Sarvam 30B and 105B, Apache 2.0) From scratch

The pattern: Southeast Asia builds sovereign models by continued pre-training on foreign open weights, increasingly Chinese ones. Korea and India train from scratch and still release open weights. Either way, sovereignty means where the model runs and who can adapt it. None of these programmes would exist on closed APIs.

The model market in three points

The model-by-model detail is in Part 2. Three points from it shape the argument.

"Open weights" does not mean "open licence". Apache 2.0 and MIT cover most of the list. The exceptions bite exactly where enterprises sit. Mistral Medium 3.5 and Devstral 2 grant no rights to companies with more than $20M monthly revenue, unless they buy a commercial licence. Qwen's largest models and Kimi K3 need a separate agreement if you resell model access at scale. MiniMax M3 bans military use. Llama 4 withholds multimodal rights from EU-domiciled companies. The NVIDIA Open Model Licence terminates automatically if you bypass its guardrails without an equivalent replacement, which matters the day you fine-tune refusals away. And licences change between versions of the same family: Gemma 4 moved to Apache 2.0, GLM-5.3 added a clause GLM-5 did not have. Review per checkpoint, not per family. Read the LICENSE file, not the model card headline.

Capability and provenance pull in opposite directions. On vendor-reported agentic coding benchmarks, the top open tier is Chinese: DeepSeek, Kimi, GLM, MiMo. The strongest non-Chinese open models (Nemotron 3 Ultra, Inkling, Mistral Large 3) sit roughly a tier below. The most transparent models are Western: Nemotron 3 Ultra publishes its training data, Apertus and Olmo publish data, code and recipes. If procurement rules out Chinese origin, you pay for it in capability. Self-hosting removes the data flow to the vendor's country. It does not change what the model was trained to do. That part is evaluation work.

The gap to closed frontier is real but narrow. On the independent Artificial Analysis Intelligence Index, the best open model scores 46 against 58 for the leading closed model and 53 for the rest of the frontier (October 2026). The gap is widest on long-horizon terminal tasks and factual recall, and narrowest on single-session coding. In practice, top open models match the closed frontier of roughly one generation ago.

For most enterprise workloads (extraction, classification, RAG over internal documents, coding agents on a known codebase), one generation behind on a model you control beats the frontier on a model you rent. Does your use case actually need the last 10 points?

Open weights still need diligence

Holding the weights moves risk, it does not remove it. Three risks move onto your side of the line: the files you download, the behaviour baked into them, and the upgrades nobody forces on you.

The files: a software supply chain

Model hubs are package registries, and they get attacked like package registries.

Date Incident What happened
Mar 2024 JFrog finds ~100 malicious models on Hugging Face Pickle payloads opened a reverse shell the moment the model loaded
Feb 2025 ReversingLabs "nullifAI" 7z-compressed PyTorch files slipped past both the loader checks and Picklescan, and still executed
Apr 2025 CVE-2025-32434 (CVSS 9.3) torch.load(weights_only=True), the standard "safe" setting, could still execute code on PyTorch 2.5.1 and earlier
2025 Three PickleScan zero-days (CVSS 9.3 each) Scanner bypasses, fixed in PickleScan 0.0.31
May 2026 HiddenLayer: fake "privacy-filter" repo Impersonated an OpenAI release, hit #1 trending with ~244k downloads in 18 hours. The weights were clean; the bundled loader.py dropped an infostealer

The last case is the lesson. Scanners check the weights. The attack came through the code shipped next to them.

Check before the first load Why
Safetensors or GGUF only; reject pickle formats No executable code at load time
trust_remote_code=False; review any bundled loader code line by line The 2026 attack vector
Official publisher repo, pinned by commit hash Download counts and likes can be faked
Verify OpenSSF Model Signing signatures where published (NVIDIA NGC signs every model) Proves file integrity and publisher, not safe behaviour
Two scanners, PyTorch 2.6 or later Single scanners have missed real payloads
First load in a sandbox without network access Contains anything the scanners missed
Serve from an internal mirror, record an AI-BOM (CycloneDX or SPDX 3.0) The inventory MAS asks for, with version and licence attached

The behaviour: what self-hosting does not fix

Origin concerns about Chinese models are mostly about data flowing to servers under PRC jurisdiction. Australia, the Czech Republic, Italy, South Korea, Taiwan and several US states acted against DeepSeek's app and service in 2025. Self-hosting solves that part. It does not solve what is inside the weights.

Concern Solved by self-hosting? Evidence
Data sent to the vendor's country Yes, with no outbound network access
App telemetry, vendor privacy policy Yes
Censorship or narrative bias No R1dacted study (May 2025): DeepSeek R1 censorship persists when run locally. NIST CAISI (Dec 2025): Kimi K2 Thinking heavily censored in Chinese, much less in English
Weak resistance to jailbreaks and agent hijacking No NIST CAISI (Sep 2025): DeepSeek R1-0528 agents 12x more likely to follow malicious instructions than US frontier models
Context-triggered weaker code No CrowdStrike (Nov 2025): politically sensitive context in prompts raised vulnerable-code rates in DeepSeek R1 by up to ~50%
Planted backdoors No, and detection is immature Anthropic, UK AISI and the Alan Turing Institute (Oct 2025): ~250 poisoned documents enough to backdoor models up to 13B
Procurement bans Read the wording Australia and the Czech Republic ban "products", which can cover self-hosted weights

Two caveats keep this honest. CAISI is a US government body comparing against US models. And the same CAISI work found DeepSeek V4 Pro roughly eight months behind the US frontier and more cost-efficient than GPT-5.4 mini on most of its benchmarks. Origin is a risk to test, not a verdict.

The practical answer is the same for every origin: put a guard model in front (Llama Guard 4, Qwen3Guard for non-English, IBM Granite Guardian 4.1 for RAG groundedness), and run your own red-team set in every language your users write in.

The upgrades: pinning cuts both ways

Closed APIs retire models on the provider's schedule.

Provider Stated minimum notice Example
OpenAI 6 months for GA models, as little as 2 weeks for previews gpt-4.5-preview: announced April 14, shut down July 14, 2025
Anthropic 60 days for public models Claude Haiku 3.5: notice December 19, 2025, retired February 19, 2026
Google Gemini Listed dates are "earliest possible" gemini-2.0-flash: about 16 months from release to shutdown

Behaviour also changes behind an unchanged model name. OpenAI rolled back a sycophantic GPT-4o update in April 2025. Anthropic's September 2025 postmortem traced degraded answers to three infrastructure bugs, at one point hitting about 16% of requests on one model.

Pinned weights do not change unless you change them. That is the control MAS asks for. The cost is that patches, security fixes and upgrades are now your release process, not your vendor's.

Self-host or API: the decision

Open weights do not mean self-hosted. The same weights are available three ways: a hosted open-weight API, a dedicated deployment in your cloud tenancy, or your own hardware. The decision is per workload, not per company.

Situation Hosted open-weight API Dedicated in your VPC Own hardware
Exploring, spiky or low volume Best fit Overkill Overkill
"Store locally" rule (copy in-country) Fine with local storage design Good fit Good fit
"Process locally" rule (inference in-country) Only with an in-country region Good fit in a local region Best fit
Regulator needs audit access to the provider Hard to negotiate Good fit Best fit
Steady high volume, predictable load Linear cost Good fit Best fit
Domain fine-tune needed If provider serves adapters Good fit Best fit
No MLOps capacity Best fit Needs a platform team Needs a platform team
Air-gapped or classified Not possible Not possible Only option

The pattern that works: start on a hosted open-weight API, build your evals against it, and move the workload in-house when one of three triggers fires. A process-locally rule, steady volume, or fine-tuning. Because the weights are identical, the move does not invalidate your evals.

With a closed model, that path does not exist. You can only renegotiate the contract.

What running it looks like

This is not theory. The rAInvent stack runs open-weight models daily, in three ways:

  • Self-hosted on own hardware: open-weight models served locally for work that should not leave the machine.
  • Hosted open-weight APIs: heavy use through Fireworks AI, Alibaba Cloud and other open-weight inference platforms.
  • Mixed agent workers: coding, review and test agents routed across DeepSeek, Kimi, gpt-oss and MiniMax, alongside Claude and Gemini. Each role gets the model that fits its cost and quality bar.

Routing like this is only possible because most of the workers are open-weight models served by more than one provider. When one provider is slow or one model underperforms on a task, the change is one line of config.

The pattern is older than LLMs. The In Mind Cloud CPQ platform was built entirely on open source, including contributions back to the Pellet OWL reasoner. The reasoning layer of a commercial product sat on a component the team could read, fix and ship without waiting for a vendor. Open-weight models bring that property to the model layer.

Where closed APIs still win

The counterargument deserves a straight answer.

  • Frontier reasoning: the strongest closed models still lead on the hardest multi-step reasoning and long-horizon agent tasks. If that gap is your use case, pay for it.
  • Safety and alignment are now yours: fine-tune aggressively and you own the validation of the resulting model. API providers carry that work for you.
  • Supply-chain diligence is now yours too: provenance, integrity checks, guard models and red-teaming are your job, and MAS names it explicitly.
  • Confidential computing is narrowing the privacy gap: NVIDIA H100 and Blackwell GPUs can run inference inside attested enclaves, with roughly 1-3% overhead on Blackwell when configured correctly. A provider that can prove it cannot read your prompts answers part of the residency objection. It does not answer exit, pinning or audit.
  • Operations are real work: GPU capacity planning, inference servers like vLLM, monitoring, upgrades. Without a platform team, self-hosting turns a model problem into an infrastructure problem.
  • Utilisation risk: a GPU node idling at 15% is more expensive per token than any API.

None of these argue against open weights. They argue against self-hosting everything. The hosted open-weight API covers most of them while keeping the exit open.

What to do with this

  • Separate the model choice from the hosting choice. Write them down as two decisions with two owners.
  • Classify each AI workload against the residency table. Store-locally, process-locally, or neither. Only the second forces local inference.
  • Run your evals on at least one open-weight model now, on your own tickets, documents or contracts, not public benchmarks. You need that baseline before a regulator or a residency rule forces the move.
  • Build the provenance check before the first download. Approved sources, checksums, licence review, a red-team pass. MAS §4.11(c) will ask for it.
  • Write the exit plan as a deployment, not a document. If your agent stack speaks an OpenAI- or Anthropic-compatible API, swapping in an open-weight model is a config change. Test that swap once a quarter and you have the exit evidence DORA, APRA and RBI ask for.

The model you pick this quarter will be outdated in two. Whether you can still run it, move it and adapt it after that depends on the decision you make now.


Part 2: Reference

The material behind the argument, for when you shortlist models and hosts.

Which models, on what hardware

Self-hosting starts with one question: what fits on hardware you can actually buy? The answer moved a long way in 2026. A single workstation GPU now runs a model that scores within reach of last year's frontier on agentic coding.

Tier Hardware Memory What fits (4-bit weights) Examples
Workstation GPU RTX 4090 / 5090 24-32 GB ~30B dense, ~35B MoE Qwen3.8-27B, Gemma 4 31B, Granite 4.2 30B, gpt-oss-20b
128 GB unified memory DGX Spark, AMD Strix Halo, Mac 128 GB 128 GB ~120B MoE gpt-oss-120b, Mistral Small 4, Nemotron 3 Super, SEA-LION v4.8 120B
512 GB Mac Studio M3 / M5 Ultra 512 GB 400-700B MoE Mistral Large 3, Nemotron 3 Ultra, GLM-5.3-Flash, MiniMax M3
One 8-GPU node 8x H100 / H200 / B200 640 GB-1.4 TB 1T-class MoE at FP8 or INT4 Kimi K2.6, DeepSeek V4 Pro, GLM-5.3, MiMo-V2.6-Pro, Inkling
Cluster Multi-node, or B300-class 1.4 TB+ 2T+ MoE Kimi K3, Qwen3.8 2.4T

The two middle tiers are single-user machines. They run a 120B or 700B model for one developer or a small team, not for hundreds of concurrent users. Production concurrency still means data-centre GPUs.

The mixture-of-experts trap still applies at every tier. "17B active" or "41B active" describes compute per token, not memory. Every parameter has to sit in memory so the router can reach any expert. Size the hardware on total parameters.

The model families, one by one

The order reflects fit for enterprise deployment (licence clarity, provenance, track record and support), not benchmark rank. On raw benchmarks, the Chinese families lead. Benchmark numbers below are vendor-reported unless marked otherwise, and every lab uses its own harness. Treat them as direction, not ranking. Hosting options list providers confirmed in the vendor documentation reviewed for this article; all of these models can also be self-hosted.

Mistral (France)

Paris-based, founded in 2023 by former Google DeepMind and Meta researchers. Europe's leading model lab and the main open-weight option for buyers who need EU origin.

  • Models: Mistral Large 3 (675B, 41B active), Mistral Small 4 (119B, 6.5B active), Ministral 3 (3B to 14B), Devstral Small 2 (24B coding), Mistral Medium 3.5 (128B dense), plus Voxtral (speech), Shieldstral (guard model) and Leanstral (formal proofs).
  • Strengths: the broadest European open line, covering text, code, speech, safety and formal verification. Most of it is Apache 2.0. Mistral Large 3 has the lowest hosted frontier-tier price in this article: $0.50 / $1.50 per million tokens. The strongest enterprise track record of any open-weight lab, with on-premises deployments at BNP Paribas, CMA CGM, Stellantis and the French armed forces.
  • Weaknesses: on coding benchmarks, Mistral's open models sit a tier below the top Chinese ones. Its strongest coders, Medium 3.5 and Devstral 2, ship under a modified MIT licence that grants no rights if "the global consolidated monthly revenue of your company (or that of your employer) exceeds $20 million". Larger companies need a commercial licence from Mistral or use the hosted API. Some products have no openly licensed weights: Mistral OCR is API-only, and Codestral weights carry the Mistral Non-Production License, so commercial self-hosting needs an agreement with Mistral. Training data is undisclosed.
  • Use cases: multilingual enterprise assistants, RAG, document work, speech, on-device, and any workload that must stay inside EU jurisdiction.
  • Where to run it: Mistral La Plateforme (EU), AWS Bedrock, Azure AI Foundry (including the EU Data Zone for Large 3 and Medium 3.5), Google Vertex, IONOS, T-Systems.

Nemotron (NVIDIA, US)

NVIDIA's own models, built to show off and sell its inference stack.

  • Models: Nemotron 3 Ultra (550B, 55B active), Nemotron 3 Super (120B, 12B active), Nemotron 3 Nano and 3.5 Lightning (30B, 3B active).
  • Strengths: the most transparent large model. NVIDIA publishes the pre- and post-training data. Ultra uses the permissive OpenMDW licence. Hybrid Mamba architecture with 1M context. NVIDIA signs every model on NGC. Ultra is the strongest US open model for reasoning and RAG. Singapore's latest SEA-LION and Malaysia's ILMU-Nemo build on Nemotron.
  • Weaknesses: Super and Nano use the NVIDIA Open Model Licence, which terminates if you bypass its guardrails without a replacement. Part of the post-training data was distilled from DeepSeek, Qwen and gpt-oss outputs, per NVIDIA's own card. The small models trail Qwen on coding.
  • Use cases: RAG, reasoning, sovereign base models, NVIDIA-standardised infrastructure.
  • Where to run it: AWS Bedrock (Super), NVIDIA NIM containers on your own GPUs.

gpt-oss (OpenAI, US)

OpenAI's first open-weight language-model release since GPT-2, published August 2025.

  • Models: gpt-oss-120b (117B, 5.1B active), gpt-oss-20b.
  • Strengths: the most widely hosted open model. Apache 2.0. The 120b fits a single 80 GB GPU or a 128 GB desktop, and is very cheap hosted ($0.15 / $0.60). Strong tool use. A popular fine-tune base (Japan's GPT-OSS-Swallow).
  • Weaknesses: no refresh in 2026. 128K context, text only.
  • Use cases: tool-calling agents, reasoning per dollar, a fine-tune base, latency-critical work.
  • Where to run it: almost anywhere. AWS Bedrock, Azure, Google Vertex, Oracle OCI, Groq (~500 tokens/s), Cerebras (~3,000 tokens/s), Together, Fireworks, OVHcloud, STACKIT.

Gemma (Google, US)

Google DeepMind's open family, small to mid-size.

  • Models: Gemma 4 31B, 26B-A4B, 12B, and the on-device E2B and E4B.
  • Strengths: moved to Apache 2.0 with version 4. Covers 140+ languages. The best on-device options in this list. The 31B fits one workstation GPU.
  • Weaknesses: the capability ceiling of a 31B model. Earlier Gemma versions (1 to 3) run under Google's own terms, which reserve a right to restrict usage, so check which version you deploy.
  • Use cases: on-device and edge, multilingual assistants, vision on a single GPU.
  • Where to run it: AWS Bedrock, STACKIT, T-Systems.

Command (Cohere, Canada)

Toronto-based and enterprise-only from the start, with a focus on private deployment.

  • Models: Command A+ (218B, 25B active).
  • Strengths: Cohere's first Apache 2.0 model. Grounded RAG with citations, 48 languages, built for air-gapped deployment. Cohere says it runs on two H100s at 4-bit.
  • Weaknesses: 128K context. Earlier Command models are non-commercial (CC-BY-NC), so check the version.
  • Use cases: on-premises RAG with citations, multilingual enterprise search.
  • Where to run it: Azure AI Foundry (including Japan, Korea, Australia and India), Cohere's own platform.

Granite (IBM, US)

IBM's enterprise-governance-first family.

  • Models: Granite 4.2 30B, 8B, 3B; Granite Guardian 4.1.
  • Strengths: Apache 2.0. Built for enterprise RAG and tool calling, with long-context retrieval. Granite Guardian checks RAG groundedness and function-call hallucinations. Fits one workstation GPU.
  • Weaknesses: not a frontier model (SWE-bench Verified 57).
  • Use cases: RAG, extraction, guardrails, regulated environments that value IBM's governance tooling.
  • Where to run it: IBM watsonx.ai, self-hosted.

Inkling (Thinking Machines, US)

The lab founded by former OpenAI CTO Mira Murati. Inkling, released July 2026, is its first large open model.

  • Models: Inkling (975B, 41B active, text, image and audio).
  • Strengths: the largest Apache 2.0 model from a US lab. Positioned as a base for customisation, with Thinking Machines' Tinker fine-tuning service.
  • Weaknesses: new. Well behind the top Chinese models on long-horizon terminal tasks (Terminal-Bench 2.1: 63.8). Node-class hardware.
  • Use cases: a US-origin base for heavy fine-tuning.
  • Where to run it: self-hosted; Tinker for fine-tuning.

Llama (Meta, US)

The model that started enterprise open-weight adoption, now in retreat. Meta's flagship Muse Spark (April 2026) is closed. The smaller Muse Glimmer (30B, August 2026) is Apache 2.0.

  • Models: Llama 4 Maverick and Scout (April 2025), Llama 3.3 70B.
  • Strengths: the longest production track record and the deepest tooling. Llama Guard 4 remains a standard guard model.
  • Weaknesses: outclassed on capability. The community licence caps use at 700M monthly users and withholds multimodal rights from EU-domiciled companies.
  • Use cases: existing deployments, guardrails, conservative environments.
  • Where to run it: AWS Bedrock, Azure, Google Vertex, Oracle OCI, Groq, OVHcloud, STACKIT, IONOS.

SEA-LION (AI Singapore)

Singapore's national model programme, part of a S$70M multimodal LLM effort.

  • Models: Nemotron-SEA-LION v4.8 (120B, 12B active, on an NVIDIA Nemotron base), Qwen-SEA-LION v4.5 27B.
  • Strengths: the best open option for Southeast Asian languages, including Burmese, Filipino, Malay, Tamil, Thai and Vietnamese. MIT licence. 262K context.
  • Weaknesses: built on foreign bases (Nemotron, Qwen), so it inherits their licence and provenance questions. Smaller tooling community.
  • Use cases: customer-facing work in Southeast Asian languages.
  • Where to run it: self-hosted.

Apertus (Swiss AI) and OLMo (AI2, US)

The fully open models: weights, training data, code and recipes all published.

  • Models: Apertus 1.5 70B (EPFL, ETH Zurich, CSCS), Olmo 3.1 32B (Allen Institute for AI).
  • Strengths: maximum transparency, which makes audits, provenance documentation and EU AI Act records straightforward. Apache 2.0.
  • Weaknesses: capability below the frontier open models.
  • Use cases: audit-heavy environments, research baselines, public sector.
  • Where to run it: self-hosted.

Qwen (Alibaba, China)

Alibaba's model family, and the most adapted open model in the world: Hugging Face counted 151,448 Qwen derivatives by August 2026, and Singapore's SEA-LION, Korea's SK Telecom A.X 4.0 and Thailand's Typhoon all build on Qwen bases.

  • Models: Qwen3.8-27B (dense, single GPU), Qwen3.6-35B-A3B (fast local MoE), Qwen3.8-Flash-Next (125B MoE), Qwen3.8 2.4T (open "Max" class).
  • Strengths: the widest size range of any family. Qwen3.8-27B is the strongest model that fits one workstation GPU, covering coding, documents and vision. Qwen3Guard is a capable multilingual guard model.
  • Weaknesses: PRC origin, which some regulated buyers rule out. Licences vary by checkpoint: the 27B is Apache 2.0, while Flash-Next and the 2.4T need a separate licence if you resell model access at scale. Alibaba also keeps some flagships API-only (Qwen3.7 Max and Plus).
  • Use cases: local coding agents, document and vision extraction, a base for regional fine-tunes.
  • Where to run it: Alibaba Model Studio (Singapore, Tokyo, Hong Kong, Frankfurt), AWS Bedrock, Together, OVHcloud, STACKIT, IONOS, Cerebras (Qwen3.8-27B).

DeepSeek (China)

A Hangzhou lab funded by the quant fund High-Flyer, known for releasing frontier-class models under MIT with no strings.

  • Models: DeepSeek V4.1 Flash (552B total, 16B active, 1M context), V4 Pro (1.6T).
  • Strengths: the top open agentic coder on vendor numbers (Terminal-Bench 2.1: 90.6). MIT licence. Low hosted prices: V4 Pro at $1.30 / $2.60 per million tokens on DeepInfra. NIST CAISI rated V4 Pro the most capable PRC model it has tested.
  • Weaknesses: the most scrutinised family. CAISI found DeepSeek R1-0528 agents far easier to hijack than US models, and censorship persists when R1 runs locally. Several governments restrict DeepSeek's apps, and some ban wording covers "products". Large hardware footprint: node-class at minimum.
  • Use cases: coding agents and long-context work where origin is acceptable and guardrails are in place.
  • Where to run it: Azure AI Foundry (V4), AWS Bedrock (V3.2), Alibaba Model Studio (V4), Together, DeepInfra.

Kimi (Moonshot AI, China)

A Beijing lab focused on long context and agentic workloads.

  • Models: Kimi K2.6 and K2.7 Code (1T total, 32B active), Kimi K3 (2.8T, 104B active).
  • Strengths: K3 leads open models on reasoning (GPQA 93.5) and document understanding. K2.6 runs on one 8x H100 node at native INT4 and is strong at agentic coding and multi-agent work.
  • Weaknesses: K3 needs a cluster. K3 ships under a custom licence that requires a separate agreement for model-as-a-service operators above $20M revenue over any consecutive 12 months. CAISI found Kimi K2 Thinking heavily censored in Chinese, much less in English.
  • Use cases: agent swarms, coding, long-document reasoning.
  • Where to run it: Azure AI Foundry (K2.5 to K2.7), AWS Bedrock, Alibaba Model Studio (K3), Together, DeepInfra.

GLM (Z.ai, China)

Z.ai, formerly Zhipu, a Tsinghua University spin-out.

  • Models: GLM-5.3 (744B, 40B active), GLM-5.3-Flash (320B, 18B active).
  • Strengths: strong at coding and cyber tasks. GLM-5.3-Flash delivers near-flagship coding at a fraction of the price, under MIT, and fits a 512 GB Mac Studio.
  • Weaknesses: GLM-5.3 added a licence clause requiring a Z.ai security review for model-as-a-service firms above $10B aggregate revenue over any consecutive 12 months. GLM-5 had no such clause, which shows why licence review has to happen per version.
  • Use cases: coding agents, security tooling, cost-sensitive reasoning.
  • Where to run it: AWS Bedrock (GLM 5), Alibaba Model Studio, Together, Fireworks, Z.ai's own API.

MiMo (Xiaomi, China)

Xiaomi's AI lab, new to frontier models but moving fast.

  • Models: MiMo-V2.6-Pro (1T, 42B active, text, image, video and audio).
  • Strengths: the highest open-model score on the independent Artificial Analysis index (46). MIT licence. Published RL training environments.
  • Weaknesses: little enterprise track record yet. Node-class hardware.
  • Use cases: general assistant and multimodal agents where capability matters most.
  • Where to run it: self-hosted; check current hosted availability.

MiniMax (China)

A Shanghai lab focused on long-context agents.

  • Models: MiniMax M3 (428B, 23B active, 1M context).
  • Strengths: efficient sparse attention for very long agent sessions.
  • Weaknesses: the most restrictive licence in this list. It requires "Built with MiniMax M3" attribution, written authorisation above $20M a year, and bans military use. The earlier M2.7 is non-commercial.
  • Use cases: long-context agent work, after legal review.
  • Where to run it: Together (M3); AWS Bedrock and Google Vertex (earlier M2 versions).

The regional long tail

  • Korea: LG K-EXAONE 2.0 (750B, Apache 2.0, Korean plus 9 languages), Upstage Solar Open2 (250B, Korean, Japanese, English), Naver HyperCLOVA X SEED, Kakao Kanana-2. Korea's national programme requires from-scratch training.
  • Japan: LLM-jp-4.1 (33B, Apache 2.0, from the National Institute of Informatics), GPT-OSS-Swallow-120B (a Japanese fine-tune of gpt-oss), Rakuten AI 3.0 (on a DeepSeek-V3 architecture).
  • India: Sarvam-105B and 30B (Apache 2.0, 22 Indian languages, trained from scratch).
  • Thailand: Typhoon 2.5 (Apache 2.0).
  • Other Chinese labs: Tencent Hy4-preview, Baidu ERNIE 4.5, ByteDance Seed-OSS-36B, Ant Group Ling-3.0, StepFun Step-3.7-Flash.
  • US niche: Arcee Trinity, Microsoft Phi-4 (small multimodal), Liquid AI LFM2 (on-device; licence capped above $10M revenue), ServiceNow Apriel.
  • UAE: TII Falcon-H1R (small hybrid models).

Best pick per use case

Use case On a workstation (up to 128 GB) On one enterprise node If origin must be outside China
Coding agent Qwen3.8-27B, Devstral Small 2 DeepSeek V4.1 Flash, GLM-5.3, Mistral Medium 3.5 (licence cap) Devstral Small 2, Mistral Small 4, Inkling
General assistant Gemma 4 31B, Mistral Small 4 Kimi K2.6, MiMo-V2.6-Pro, Mistral Large 3 Mistral Large 3, Command A+, Nemotron 3 Ultra
RAG and extraction Granite 4.2 30B, Qwen3.8-27B, Mistral Small 4 Command A+, Mistral Large 3 Command A+, Mistral Large 3, Granite, Nemotron 3 Ultra
Southeast Asian languages Qwen-SEA-LION v4.5 27B Nemotron-SEA-LION v4.8 SEA-LION v4.8, Gemma 4
Documents and vision Qwen3.8-27B, Gemma 4, Mistral Small 4 Kimi K3, Mistral Large 3 Gemma 4, Mistral Small 4, Mistral Large 3
Speech Voxtral Mini 3B, Voxtral Realtime 4B Voxtral Small 24B Voxtral family
Guardrails Shieldstral 1.0, Qwen3Guard, Granite Guardian 4.1 Llama Guard 4 Shieldstral, Granite Guardian, Llama Guard
Formal verification Leanstral 1.5 Leanstral 1.5 Leanstral
On-device Gemma 4 E2B / E4B, Ministral 3 n/a Gemma 4, Ministral 3, Phi-4
Maximum transparency Apertus 1.5, Olmo 3.1 Nemotron 3 Ultra All three

Among European labs, Mistral is the only one with an open model in most rows, including speech and formal proofs. Its Apache 2.0 line (Large 3, Small 4, Ministral 3, Devstral Small 2, Voxtral, Shieldstral, Leanstral) is the default shortlist for buyers who need EU origin. One exception inside the family: Voxtral TTS (text-to-speech) is non-commercial.

Where to rent open weights

"Start on a hosted open-weight API" raises the next question: whose? Model catalogues, regions and data defaults differ more than the marketing suggests. Status per vendor documentation, October 3, 2026.

Vendor Jurisdiction Type Open models (selection) APAC regions EU regions Data defaults
AWS Bedrock US Hyperscaler gpt-oss, DeepSeek V3.2, Qwen3, Kimi K3, GLM 5, MiniMax M2.5, Mistral Large 3, Llama 4, Gemma 4, Nemotron 3 Super In-region in Jakarta, Tokyo, Mumbai, Sydney, Melbourne Frankfurt, Stockholm, Milan, Ireland, London (per model) Zero retention configurable per region, enforceable org-wide. Model vendors see no prompts
Azure AI Foundry US Hyperscaler DeepSeek V4, Kimi K2.5-K2.7, Llama 4, Mistral Large 3, gpt-oss-120b Australia East, Japan East/West, Korea Central, South India; global deployment, no APAC data zone EU Data Zone for DeepSeek V4-Flash, Mistral Large 3, Mistral Medium 3.5 Global deployments may process in any Azure region
Google Vertex AI US Hyperscaler DeepSeek V3.2, Qwen3, gpt-oss, Llama 4, Kimi K2, MiniMax M2, GLM 5.2, Mistral Processing: gpt-oss and Qwen3 in the US, DeepSeek V3.2 global. No Asian region Same: US or global processing DeepSeek V3.2 and Qwen3-235B endpoints retire October 21, 2026
Oracle OCI US Hyperscaler Llama 4, Llama 3.3, gpt-oss Osaka on-demand, Hyderabad dedicated Frankfurt Dedicated clusters available
Alibaba Cloud Model Studio China (parent) APAC cloud Qwen 3.x, DeepSeek V4, Kimi K3, GLM 5.3 Singapore, Tokyo, Hong Kong Frankfurt No training on customer data. Inference can route globally unless you pick a geo-bounded scope
Fireworks AI US Inference specialist Broad open catalogue Dedicated: Malaysia, Tokyo, New South Wales Dedicated: Frankfurt, Iceland Zero retention by default for open models
Together AI US Inference specialist DeepSeek V4.1, Kimi K3, GLM-5.3, MiniMax M3, Qwen 3.x, gpt-oss No region choice on serverless Via dedicated endpoints or VPC Stores prompts by default; zero retention is opt-in
DeepInfra US Inference specialist Broad open catalogue Not stated Not stated In-memory only, no training
Groq US Speed specialist (~500 tokens/s on gpt-oss-120b) gpt-oss, Llama 3.x, previews None None No retention by default; stored data sits in US buckets
Cerebras US Speed specialist (~3,000 tokens/s on gpt-oss-120b) gpt-oss-120b, Qwen3.8-27B Not stated Not stated SOC 2 Type 2
Baseten US Inference platform Any model, custom weights Region restrictions configurable Region restrictions configurable Zero retention by default; can run inside your own VPC. SOC 2 Type II, HIPAA
OpenRouter US Aggregator Everything, routed to other hosts Depends on upstream host EU in-region routing on business plans Per upstream host; can filter out hosts that train on data
Hugging Face Inference Providers US / France Router Routes to Fireworks, Groq, Together, OVHcloud, Scaleway and others Depends on provider Pin EU providers Provider picked dynamically unless pinned
Mistral La Plateforme France EU sovereign Mistral Large 3, Small 4, Ministral 3 n/a EU Enterprise accounts opted out of training; zero retention available
OVHcloud AI Endpoints France EU sovereign Llama 3.3, Qwen 3.x, gpt-oss n/a EU Not stated on catalogue
IONOS AI Model Hub Germany EU sovereign Llama, Mistral, Qwen 3.5 n/a Germany Not stated on catalogue
STACKIT AI Model Serving Germany (Schwarz Group) EU sovereign gpt-oss, Llama 3.3, Qwen3-VL, Qwen3.8-27B, Gemma n/a EU01 Not stated on catalogue
T-Systems AI Foundation Services Germany EU sovereign Mistral, Qwen, Gemma n/a T-Cloud Germany; some models served via Azure Sweden or Google Cloud ISO 27001

Also in the market, not covered here: Scaleway, Nebius, Aleph Alpha, BytePlus ModelArk, Huawei Cloud, Tencent Cloud, Baidu Qianfan, and local GPU clouds such as Singtel RE:AI, Sakura Internet, Naver Cloud and Yotta, which mostly sell GPU capacity rather than managed model APIs.

Six things the table shows.

Check where inference actually runs, not where the endpoint lives. Region lists tell you where you can create an endpoint. They do not always tell you where your data is processed. AWS Bedrock runs its open models in-region, in Jakarta, Tokyo, Mumbai and Sydney among others. Azure offers them as global deployments that may process in any Azure region. Google processes gpt-oss and Qwen3 in the US and DeepSeek V3.2 globally. Alibaba Cloud serves Qwen, DeepSeek, Kimi and GLM from Singapore, Tokyo, Hong Kong and Frankfurt, and needs a geo-bounded scope to keep inference in-region. If a customer needs processing in a specific country, choose the host by its processing location, or deploy the weights yourself on GPU instances in that region.

Model origin and data jurisdiction are separate choices. You can run DeepSeek, Qwen or Kimi on AWS in Frankfurt, or Qwen and gpt-oss on STACKIT in Germany. AWS documents that model providers have no access to prompts or completions. With open weights, the origin risk stays in the weights and the licence. It leaves the data path.

Every host carries some jurisdiction. US-headquartered hosts fall under the US CLOUD Act, including the specialists and routers. Alibaba swaps that for PRC-law exposure through its parent. The EU-headquartered options (Mistral, OVHcloud, IONOS, STACKIT, T-Systems for its German-hosted models) sit outside both. For French sensitive state data under SecNumCloud, that difference decides eligibility.

Retention defaults differ. Fireworks, DeepInfra, Baseten and Groq keep nothing by default. Together stores prompts unless you opt out. Read the data-handling page before the pricing page.

Hosted catalogues churn too. Bedrock commits to keeping a model at least 12 months after launch. Google deprecated its DeepSeek V3.2 and Qwen3-235B endpoints in July 2026, with retirement on October 21, three months later. Specialist serverless catalogues turn over faster still, and preview tiers are evaluation-only. The pinning argument applies to hosted open models as well: if a workload must not change, use a dedicated deployment or your own weights. The difference from a closed API is that the exit stays open when the host drops the model.

Hyperscaler or specialist is a trade. Hyperscalers bring existing contracts, IAM, private networking and committed spend, but lag on new models: Bedrock serves DeepSeek V3.2 while Azure and Alibaba already offer V4. Specialists ship new models within days, are cheaper, and offer more dedicated regions, but publish thinner compliance evidence. Bring-your-own fine-tunes are geographically narrow on hyperscalers (AWS Custom Model Import runs in Frankfurt and three US regions; Azure open-model fine-tuning is US-only).


Sources

Models and hosting

Economics and adoption

Hosting providers (all accessed October 3, 2026)

Security and control

Singapore

EU

APAC residency

Read more

GenAI Daily - October 3, 2026: OpenAI Fires Safety Researchers, Decision Models Go Open-Weight, ServiceNow Launches Flow

GenAI Daily - October 3, 2026: OpenAI Fires Safety Researchers, Decision Models Go Open-Weight, ServiceNow Launches Flow

Top Stories OpenAI Fires Three Safety Researchers Over Information Shared With Outside Evaluators OpenAI confirmed it parted ways with Jasmine Wang, Tomek Korbak and Mikita Balesni for violating its policies on accessing and handling sensitive company information. The Wall Street Journal first reported that the alleged misconduct included sharing confidential

By Falk Brauer