Skip to main content

AI decisions Architecture and technology

Sovereign or on-prem AI vs cloud AI APIs: where should your models run?

Short answer

Decide by data, not by ideology. Cloud AI APIs are the default for most workloads because they give strong models with little to operate; sovereign or on-prem deployment is justified when the data, the regulator or the contract requires that processing stays under your control and EU jurisdiction. Most European organisations end up hybrid: sensitive workloads on infrastructure they control, the rest on cloud APIs under a governed gateway.

Updated: · Thinkia

The options

Sovereign / on-prem AI

Models and retrieval run on your own data centre, a private cloud or a provider under EU jurisdiction that you control contractually and technically.

Cloud AI APIs

Models consumed as a managed service from a hyperscaler or model provider, possibly in EU regions, under that provider's terms.

Side by side

Criterion Sovereign / on-prem AICloud AI APIs
Where data is processed Inside a perimeter you define and can audit. In the provider's infrastructure; EU regions are often available, depending on the service.
Jurisdiction Can be kept fully under EU law if the operator is also EU-based. A non-EU provider may be subject to third-country laws with extraterritorial reach, even with data stored in the EU.
GDPR international transfers Transfers can be avoided by design. Requires a valid transfer basis (adequacy, standard contractual clauses) where personal data may leave the EEA, and a transfer impact review.
Model choice Mostly open-weight models you can host; frontier proprietary models are rarely available this way. Access to the strongest proprietary models as soon as they are released.
Operational burden Hardware or capacity, MLOps, security patching and monitoring are yours. Largely managed by the provider.
Cost profile Upfront and fixed; pays off with steady volume and long horizons. Pay per use; flexible, but can be hard to predict at scale.
Auditability Full logs and change control under your rules. Depends on what the provider logs, exposes and contracts.
EU AI Act Deployer obligations are the same; location does not reduce them. Same deployer obligations; you also depend on the provider's documentation as GPAI or system provider.

Choose Sovereign / on-prem AI when…

  • You process special categories of personal data, classified information or core intellectual property that policy says cannot leave your control.
  • You are a public body or critical operator with security frameworks that constrain where and by whom data may be processed (in Spain, for example, the Esquema Nacional de Seguridad).
  • A sector regulator, auditor or contract requires full traceability of where inference runs and which model version answered.
  • The workload is steady and high-volume, and a well-tuned open-weight model meets the quality bar.

Choose Cloud AI APIs when…

  • The data is public, internal-low-sensitivity or can be anonymised before it leaves your perimeter.
  • You need frontier capability that is only offered as a service.
  • Demand is uncertain and you want to validate value before investing in infrastructure.
  • Your team cannot yet operate models in production securely.

When to combine them

A hybrid pattern works for most organisations. Classify each workload by data sensitivity and regulatory impact; keep retrieval over sensitive documents and the most sensitive inference on infrastructure you control; anonymise or minimise data locally before any call to an external model; and send the rest to cloud APIs through a single gateway that enforces identity, logging and routing rules. The rule should live in the gateway, not in each team's code.

Common mistakes

  • Confusing data residency with sovereignty: an EU region does not by itself remove third-country jurisdiction over the provider.
  • Going on-prem for everything and ending up with weaker models, high fixed cost and slow delivery.
  • Sending sensitive data to a public API during a pilot and discovering the legal blocker only before production.
  • Assuming that running a model locally exempts you from the EU AI Act; deployer duties depend on the use case.
  • Leaving security, legal and the data protection officer out of the architecture until the end.

How Thinkia approaches it

We start with a classification, not a platform. Each use case is placed by data sensitivity, regulatory exposure and required capability, and that decides where it runs. Security, legal and the data protection officer join the architecture from the first session, because they are the ones who will approve or block production.

Synapse is built for this hybrid reality. Its RAG repository runs on the client's own servers, so documents and embeddings do not leave the corporate perimeter; its gateway enforces corporate SSO and logs every call; and its routing prioritises confidential local models before commercial endpoints. Because it is LLM-agnostic, a workload can move between a hosted open-weight model and a cloud API without rewriting the application.

On regulation, we treat sovereignty and compliance as separate questions. GDPR governs where personal data goes and on what legal basis; the EU AI Act (Regulation (EU) 2024/1689) governs what the system does and who is accountable. A sovereign deployment can still be high-risk, and a cloud deployment can be compliant. For public bodies we also consider national frameworks and the supervisory authority, in Spain AESIA. This is practical orientation, not legal advice.

Thinkia products involved

Related AI solutions

Frequently asked questions

Is hosting in an EU cloud region enough to be “sovereign”?

Not always. Residency means the data is stored and processed in the EU; sovereignty also covers who can be compelled to access it and under which law. If the provider is subject to non-EU legislation with extraterritorial reach, assess that risk explicitly, together with contractual and technical safeguards such as encryption with keys you control.

Can we use cloud AI APIs with personal data under GDPR?

Yes, if you have a legal basis, a data processing agreement, appropriate safeguards for any transfer outside the EEA and, where required, a data protection impact assessment. Transfer frameworks have been struck down before, so keep a fallback option for the most sensitive workloads.

Does on-prem AI mean worse models?

Often it means different models. The strongest proprietary models are rarely available to host yourself, but open-weight and smaller specialised models perform well on many bounded tasks. Test on your own cases before deciding the quality gap matters.

What does the EU AI Act change for this decision?

It does not dictate where models run. It sets obligations by risk level and role: as deployer you need human oversight, logging and relevant input data for high-risk uses, wherever the model is hosted. Annex III high-risk obligations apply from December 2027, after the Digital Omnibus on AI entered into force on 27 July 2026; check the consolidated text on EUR-Lex or the AI Act Service Desk. This is not legal advice.

Is sovereign AI mandatory for the public sector?

There is no single rule. It depends on the data, the security category of the system, national frameworks and procurement conditions. Many public bodies combine both: citizen-facing information services on cloud APIs and case files or sensitive records on controlled infrastructure.

Where should we start?

Inventory current and planned AI use cases, including unsanctioned ones, and classify them by data sensitivity. That map usually shows that only a minority of workloads need sovereign hosting, and those are the ones to design first.

Related decisions

Sectors where this decision comes up

Key terms

Thinkia articles

Whitepapers

Sources

Facing this decision now? Talk it through with us.

Talk to an AI Expert