Every European LLM API provider reviewed
Sweden
Founded 2024
Trial €0 with €5 starting credits / Starter €25 per month / Team €50 per month / Dedicated endpoints via sales
Trial plan with €5 of starting credits
Best for: Teams that want a small European inference provider with a stated zero-retention position and a flat monthly plan
Berget AI is a Swedish inference company founded in 2024. Its home page offers serverless inference through an OpenAI-compatible API, dedicated endpoints for custom or fine-tuned models, and add-ons such as vision, audio understanding, OCR, speech transcription, embeddings and reranking. The code sample points the OpenAI SDK at api.berget.ai and calls gemma-4-31B-it.
The privacy policy names Berget AI AB, organisation number 559504-7522, as data controller.
The home page states that it never stores the content of prompts or generated output and never trains on customer data, and that the DPA is included with all plans. The DPA lists the sub-processors used: Swedish affiliates for network operation, GPU bare-metal infrastructure and colocation, and Stripe Payments Europe for card payments. The pages read say models run under Swedish jurisdiction and data stays in the EU, but name no city.
Pricing is published as plans: a Trial at €0 with €5 of starting credits, Starter at €25 per month and Team at €50 per month with higher rate limits and priority support, billed as a monthly top-up that is consumed by usage. Dedicated endpoints go through sales.
What Berget AI does well
- Zero data retention stated on the home page, with a DPA on every plan
- Sub-processors named in the DPA, all Swedish affiliates plus Stripe
- Plans published from €25 per month, with a free trial credit
- OpenAI-compatible API plus embeddings, reranking, OCR and transcription
Where Berget AI falls short
- No data centre city named on the pages read
- The privacy policy read is dated March 2025
- Rate limits on the low plans are described only as basic or increased
- No security certification named on the pages read
Standout feature. It is the only provider here that pairs a zero-retention statement on its home page with a DPA that lists named sub-processors by organisation number.
Geneva, Switzerland
Pay per use in CHF, for example Gemma 4 31B at CHF 0.20 per 1M input and CHF 0.40 per 1M output tokens / Kimi K2.6 at CHF 0.60 and CHF 3.00 / 1M free credits
1M free credits
Best for: Teams that accept Switzerland as the location and want a published per-model price list
Infomaniak is a Geneva cloud company. Its AI Services page offers open-source models through OpenAI-compatible APIs, with LLM assistants, semantic search, function calling, audio transcription and image generation, and states that its AI and GPU services run in its data centres in Switzerland with no dependence on foreign players.
The legal notice names Infomaniak Network SA at Rue Eugène Marziano 25 in Les Acacias, Geneva, with Commercial Register number CH-660.0.059.996-1. The AI Services page says no queries are stored and that the energy for the AI infrastructure is renewable and recovered for heating.
The price list is published in Swiss francs per million tokens. Gemma 4 31B costs CHF 0.20 input and CHF 0.40 output, Mistral Small 4 CHF 0.20 and CHF 0.75, Kimi K2.6 CHF 0.60 and CHF 3.00, and Qwen3.5-397B CHF 0.80 and CHF 3.60. Embeddings start at CHF 0.065 per million tokens, Whisper V3 is CHF 0.006 per minute, and new accounts get 1M free credits.
What Infomaniak does well
- Inference stated to run in data centres in Switzerland
- No queries stored, per the AI Services page
- Full price list in Swiss francs, including embeddings, audio and image models
- 1M free credits for new accounts
Where Infomaniak falls short
- Switzerland is outside the EU, covered by an EU adequacy decision
- Prices are in Swiss francs, which adds exchange-rate exposure for a euro budget
- The terms and DPA were not read in full for this review
Standout feature. It is the only provider here that publishes a complete price list for language, embedding, reranking, audio and image models on one page.
Paris, France
Per million tokens in euro, for example gemma-4-26b-a4b-it at €0.25 input and €0.50 output / 1,000,000 free tokens for new customers / Dedicated Deployment priced separately
1,000,000 free tokens for new customers
Best for: Developers who want OpenAI-compatible calls served from one named French data centre
Scaleway is a Paris cloud provider whose Generative APIs serve open and Mistral models through a serverless API, with Dedicated Deployment for customers who need reserved capacity. The documentation has a page on OpenAI API compatibility, and the model list covers chat, code, vision, embeddings and speech.
The legal notice names SCALEWAY SAS, SIREN 433 115 904, RCS Paris, at 8 rue de la Ville-l'Évêque in Paris. The FAQ states that all models are hosted in a secure data centre in Paris and that this may change in the future. The product page states that Scaleway does not collect, read, reuse or analyse the content of inputs, prompts or outputs.
Prices are published in euro per million tokens, for example €0.25 input and €0.50 output for gemma-4-26b-a4b-it and €1.50 and €7.50 for mistral-medium-3.5-128b, with cached-token prices on some models. New customers get 1,000,000 free tokens before paying from the next token.
What Scaleway does well
- Inference stated to run in one data centre in Paris
- States that prompt and output content is not collected or reused
- Euro prices per model, with cached-token rates and 1,000,000 free tokens
- Dedicated Deployment for reserved capacity
Where Scaleway falls short
- The FAQ says the Paris-only hosting may change in the future
- Models are open-weight and Mistral, not GPT or Claude
- Output limits on serverless endpoints are capped to prevent runaway bills
Standout feature. It is the only provider here whose own FAQ names a single city for all model hosting.
Montabaur, Germany
Per million tokens in US dollars, for example gpt-oss-120b at $0.17 input and $0.71 output / Qwen3.5-397B-A17B at $0.67 and $4.00
Best for: Buyers who want the strictest stated data path from a large German provider
IONOS AI Model Hub is a managed model API inside IONOS Cloud, with large language models, coding models, text-to-image, OCR, embedding and reranking models. The documentation describes an OpenAI-compatible REST API, and the example endpoint is labelled de-txl.
The imprint names IONOS SE at Elgendorfer Str. 57, 56410 Montabaur, registered at Amtsgericht Montabaur under HRB 24498.
The data handling page states that all inference is performed in IONOS data centres in Germany, that prompts and outputs are never logged or written to persistent storage, that the API is stateless per request, that operational access is limited to personnel inside the EU, and that no third-party sub-processors are engaged. Only call metadata (timestamp, model, token count) is kept, for billing.
Pricing is per million tokens in US dollars: Qwen3.5-9B at $0.11 input and $0.17 output, gpt-oss-120b at $0.17 and $0.71, Llama 3.3 70B at $0.71 and $0.71, and Qwen3.5-397B-A17B at $0.67 and $4.00.
What IONOS does well
- All inference stated to run in German data centres
- Prompts and outputs never logged or stored, per the data handling page
- No third-party sub-processors and EU-only operational access
- Per-model price list and OpenAI-compatible API
Where IONOS falls short
- Prices are listed in US dollars
- The catalog turns over: models such as Llama 3.1 8B and Mistral Nemo are scheduled for retirement in October 2026
- A large vendor with a broad cloud catalog, so the AI product is one entry among many
Standout feature. It is the only provider here whose documentation states that no third-party sub-processors are engaged for the model API.
Paris, France
Per million tokens, with Mistral Large given as $0.5 input and $1.5 output / Batch 50% lower / cached input up to 90% lower
Best for: Teams that want a European model maker with its own API and an explicit EU or US endpoint choice
Mistral AI is a Paris company that builds its own models and sells them through Studio and an API, next to coding and enterprise products. The pricing page says most models are priced per million tokens, with input and output counted separately and Mistral Large given as an example at $0.5 input and $1.5 output. Batch processing is 50% lower and cached input up to 90% lower.
The privacy policy names Mistral AI as a French company incorporated in Paris under number 952 418 325, at 15 rue des Halles. The Help Center says data is hosted in the European Union by default, that a US API endpoint is optional, and that data can be temporarily transferred outside the EU to the sub-processors listed in its Trust Center, with Standard Contractual Clauses.
Zero Data Retention is available only with pay-as-you-go and only for stateless endpoints such as chat completions, not for conversations, agents, libraries or batch files. Without it, the Help Center says data is stored for as long as it is processed, with account, technical and invoice records kept for the legal periods it lists.
What Mistral AI does well
- A European model maker with its own models rather than only hosted ones
- EU hosting by default, US endpoint only on request
- Zero Data Retention available on pay-as-you-go for stateless calls
- Batch and cached-token discounts published
Where Mistral AI falls short
- Data can be transferred outside the EU to listed sub-processors
- Zero Data Retention excludes stateful products and the free tier
- Prices were read in US dollars on the page, with euro shown on selection
- No OpenAI-compatibility claim was found on the pages read
Standout feature. It is the only provider here that is both the model maker and the host, so the model weights and the API come from one company.
Roubaix, France
Per million tokens in euro, for example gpt-oss-120b at €0.08 input and €0.40 output / Qwen3.5-397B-A17B at €0.60 and €3.60 / Batch 50% off
Playground for free; Public Cloud free trial credit advertised on the same site
Best for: Teams that want a large French provider with a zero-retention statement and an OpenAI-compatible API
OVHcloud AI Endpoints is a serverless inference API for more than 20 models such as Llama, Qwen and gpt-oss, plus embeddings, speech and a guard model. The product page lists OpenAI-compatible standard APIs, a Playground to test models for free, batch processing at 50% off, integrations with Hugging Face, LiteLLM, Continue and Kilo Code, and a model catalog with prices in euro.
The pages carry the OVH SAS copyright line and the OVH Groupe SA entity at 2 rue Kellermann, 59100 Roubaix.
The product page says the platform is hosted by OVHcloud, a European provider of sovereign cloud services since 1999, that data is never used to train its models, and that it keeps only the data required for billing. The documentation says its infrastructure is located in Gravelines, France, and the product page feature table adds "Worldwide availability (non-EU)" beside the SLA, which is unclear on the pages read, so confirm in writing where inference runs for your account.
Catalog prices are per million tokens: gpt-oss-120b at €0.08 input and €0.40 output, Qwen3.5-397B-A17B at €0.60 and €3.60, Qwen3.8-27B at €0.40 and €2.70, and Qwen3-Embedding-8B at €0.10 input.
What OVHcloud does well
- Zero data retention except billing data, per the product page
- OpenAI-compatible API with a free Playground
- Euro catalog prices and a 50% batch discount
- Models can be deployed on the customer's infrastructure to avoid lock-in, per the product page
Where OVHcloud falls short
- The "Worldwide availability (non-EU)" line in the feature table is unexplained
- The base tier lists a 99.5% to 99.98% SLA in different places
Standout feature. It is the only provider here that lists a free Playground for every model next to a batch tier at half price.
Vienna, Austria
Pay as you go at the provider price plus a 5% flat service fee, billed in euro by card or invoice
Best for: Teams that want one European gateway that can be limited to EU-established providers
Cortecs is a Vienna company that sells an LLM router: one API, one invoice, and policy-based routing across providers. Its home page lists 16 active providers, including Berget, IONOS, Scaleway, OVH and Nebius, and also Amazon, Google and Microsoft, supports OpenAI or Anthropic endpoints, and adds budgets, usage tracking per team or key, and a model catalog of more than 150 models.
The imprint names Cortecs GmbH at Althanstraße 4 in Vienna, company register number 560802i, Commercial Court of Vienna, and the terms apply Austrian law with Vienna courts. The DPA states that payload data is processed only in volatile memory and deleted on completion of the routing request, that operational metadata may be kept for up to twelve months, and that data sent to an inference sub-processor follows that provider's own retention policy.
Pricing is pay as you go at the provider's direct price plus a 5% flat service fee, in euro by card or invoice. The Sovereign Cloud option restricts routing to providers established in the EU.
What Cortecs does well
- Sovereign Cloud mode limits routing to EU-established providers
- One API and one euro invoice across many providers
- Payloads held only in volatile memory, per the DPA
- Budgets and usage tracking by team, user, key and provider
Where Cortecs falls short
- Outside Sovereign Cloud mode, requests can reach Amazon, Google or Microsoft
- Retention at the end provider depends on that provider
- A gateway fee of 5% is added to every provider price
Standout feature. It is the only provider here whose DPA states that retention at the sub-processor is the customer's responsibility to configure.
Denmark
€20 per month, up to 4,500 requests in a rolling 5-hour window
Best for: Developers who want a flat monthly price for a coding agent on EU-hosted inference
2BA.AI is a Danish service, according to the vendor, that sells an OpenAI-compatible API aimed at coding agents. Its home page describes an installer that pairs the browser, creates an API key and configures the coding tool, with a plan of €20 per month for up to 4,500 requests in a rolling 5-hour window. It publishes an internal evaluation on 50 tasks from SWE-bench Verified, where its own "Amber" model at extra-high effort resolved 44 of 50.
The privacy policy states that API inference, including processing of inputs and outputs, is hosted within the European Union, that request content is processed transiently for inference and not stored as prompt or completion history, and that inputs and outputs are not used to train or improve models. Account, billing and operational metadata are separate from request content, Stripe processes subscription payments, and erasure requests go to privacy@2ba.ai.
The terms and privacy policy do not name the company or its address: they define "we" as the service provider identified in the order or billing documentation. The Danish establishment therefore rests on what the vendor told us and should be checked against the order form. The terms say renewal can be cancelled by either side at any time, that fees for a started period are not refunded, and that there is no uptime guarantee unless agreed in writing.
What 2BA.AI does well
- Flat €20 per month with a stated request allowance per 5 hours
- EU-hosted inference and no prompt history, per the privacy policy
- OpenAI-compatible, with an installer for common coding tools
- Referral terms and erasure contact published
Where 2BA.AI falls short
- The terms and privacy policy name no legal entity or address
- The 50-task benchmark is the vendor's own and small
- No SLA, uptime or model-availability guarantee in the terms
- The vendor can decline to renew a subscription at any time
Standout feature. It is the only provider here that sells access as a request allowance per rolling 5-hour window instead of per token.
Bad Friedrichshall, Germany
Best for: Companies already on a Schwarz Group cloud that want a model API with no content stored
STACKIT AI Model Serving is a managed hosting environment for language models such as Llama and Mistral inside the STACKIT cloud. The documentation describes an inference API that is OpenAI compatible, auth tokens, rate limits, tool calling, vision, and tutorials for LangChain, retrieval-augmented generation and OpenCode.
The site footer names Schwarz Digits Cloud GmbH & Co. KG at Am Campus 1 in Bad Friedrichshall. The FAQ states that no customer data from requests is stored or used by STACKIT, that no model is trained on customer data, and that only open-source models are served. The AI Model Serving pricing table names the region Germany South, and the imprint page itself returned no content.
The pages read do not publish a token price, and point to the STACKIT calculator for costs, which could not be read as it loads in the browser.
What STACKIT does well
- No customer data from requests stored or used for training, per the FAQ
- OpenAI-compatible inference API with tool calling and vision
- Sits inside a full European cloud catalog
Where STACKIT falls short
- No token price on the pages read
- The shared-model list is selected by STACKIT, with manual instance selection not possible
- The imprint page returned no content, so the entity comes from the footer
Standout feature. It is the only provider here whose FAQ says plainly that the model is assigned by availability and cannot be picked by instance.
Amsterdam, Netherlands
Best for: Teams that need large open-model capacity and will pay for dedicated endpoints in a chosen region
Nebius Token Factory, the successor to Nebius AI Studio, is an inference API for open models with an OpenAI-compatible API, dedicated GPUs, post-training and workload optimisation. The documentation covers function calling, structured output, rate limits, observability and fine-tuning.
The site footer names Nebius B.V., and the legal guide says the contracting entity is the applicable Nebius entity named in the terms, part of Nebius Group N.V., which is listed on NASDAQ.
The guide states that public endpoints show the region Global, that processing can move at any time without notice, and that dedicated endpoints stay in the chosen region as a contractual commitment under the DPA. Prompts and outputs may be stored in Finland for speculative decoding unless Zero Data Retention is enabled, which also turns off that feature and may slow inference.
The pages read give no price list, only a usage-based model, so the cost needs to be taken from its price page or a sales call.
What Nebius AI Studio does well
- Dedicated endpoints stay in the chosen region, a contractual commitment under its DPA
- Zero Data Retention switchable per organisation, with no training either way
- Fine-tuning and custom model weights supported
- The guide states plainly where data is stored when ZDR is off
Where Nebius AI Studio falls short
- Public endpoints have no region commitment and are called unsuitable for compliance-driven production
- Zero Data Retention is off by default and may reduce speed
- The contracting entity depends on the customer and is named only in the terms
- No price list on the pages read
Standout feature. It is the only provider here whose own documentation tells customers not to run region-sensitive production traffic on its default endpoints.
Bratislava, Slovakia
Per million tokens in US dollars, for example gpt-oss-20b at $0.05 input and $0.20 output / no minimum spend or subscription
Best for: Teams that want one EU contract and one key across models on own EEA hardware and routed models
Heabsy is a Bratislava company with an AI platform for regulated industries, including an inference API that is OpenAI- and Anthropic-compatible. Its model catalog lists 31 models with per-million-token prices, context length and capabilities, and sorts them into two tiers: models on GPUs it operates inside the EEA with zero data retention, and routed models fulfilled through global providers under its European contract.
The legal notice names FEYA, s.r.o., brand Heabsy, at Pekná cesta 19 in Bratislava, registered at Mestský súd Bratislava III, file no. 100404/B, IČO 47 887 541. The privacy policy says that for content sent to the inference API Heabsy acts as processor under a DPA. The catalog says routed compute may run outside the EEA and says so per model instead of hiding it; the terms say some models may be fulfilled through third-party inference providers.
Prices are per million tokens in US dollars with no minimum spend and no subscription, for example gpt-oss-20b at $0.05 input and $0.20 output and DeepSeek-V4-Flash at $0.06 and $0.12. The consulting side of the business starts at €5K per project.
What Heabsy does well
- A catalog that marks each model as own EEA hardware or routed
- Zero retention stated for the own-hardware tier
- Per-model prices with no minimum spend
- OpenAI- and Anthropic-compatible endpoints under one DPA
Where Heabsy falls short
- Many of the listed models are marked as processed outside the EEA
- Prices are in US dollars
- SOC 2 is on a roadmap, not held
- The inference product sits beside a consulting business whose projects start at €5K
Standout feature. It is the only provider here that labels every model in its catalog with where the compute runs, including the ones that run outside the EEA.
Utrecht, Netherlands
API needs a Pro, Teams or API-only subscription / chat plans from €4.50 to €17.50 per month / DeepSeek V4.1 Flash at €0.22 input and €1.10 output per 1M tokens / glm-5.3 at €1.10 and €4.40
14-day free trial, no credit card
Best for: Teams that want a small Dutch provider on French hosting and a measured energy figure
GreenPT is a Utrecht company that runs a chat product and an API. The API is OpenAI-compatible, with chat completions, streaming, embeddings, reranking, OCR, speech-to-text, scraping and web search through one key, and a Green Router that picks an efficient model per request. The catalog lists open-weight models such as DeepSeek V4.1 Flash, glm-5.3, kimi-k3 and MiniMax.
The terms and privacy policy name GreenPT BV at Plompetorengracht 4 in Utrecht, Chamber of Commerce 97084360, under Dutch law with Amsterdam courts.
The privacy policy states that personal data is processed in EU data centres in France and that API payloads are processed in memory only and not stored. Its sub-processor table names Scaleway SAS in France for hosting, Inceptron, Tensorix and Melious on Verda data centres in Finland for AI hosting, and Mollie B.V. for payments. The privacy policy is dated January 2026.
API use requires an active Pro, Teams or API-only subscription. Chat plans run from €4.50 to €17.50 per month, with a 14-day free trial, and models are priced per million tokens in euro, for example €0.22 input and €1.10 output for DeepSeek V4.1 Flash and €1.10 and €4.40 for glm-5.3.
What GreenPT does well
- Hosting provider and AI sub-processors named with country
- API payloads in memory only and not stored, per the privacy policy
- Euro prices per model and a 14-day trial without a card
- One key for chat, embeddings, reranking, OCR, speech and search
Where GreenPT falls short
- The API requires a paid subscription on top of token prices
- The privacy policy names Finland as well as France for AI hosting
- ISO 27001 is claimed on the site without a certificate shown on the pages read
- The sub-processor list includes another provider in this ranking, Melious AI
Standout feature. It is the only provider here that shows an energy and CO2 figure per chat and names its data centre efficiency figures with a source.
Saarbrücken, Germany
Pay as you go, for example DeepSeek V4.1 Flash at €0.20 input and €1 output per 1M tokens / coding plans Basic €49, Pro €149, Unlimited 3x €499, Unlimited 6x €999 per month, businesses only, plus VAT
Best for: Businesses that want one German contract for a gateway across European inference providers
Melious AI is a Saarbrücken company that offers inference on more than 60 open-weight models through an OpenAI- and Anthropic-compatible API, plus web search, URL fetching, code execution and managed vector storage. The home page lists 11 providers in 8 countries and says an independent check by staysin.eu covers its web app and API. The plans page sells coding plans and pay-as-you-go credits.
The legal notice names Melious AI GmbH at Universität des Saarlandes, Campus Starterzentrum, Saarbrücken, registered at Amtsgericht Saarbrücken under HRB 111794, with Simon Jakobs as managing director.
The DPA lists the sub-processors, among them IONOS SE in Germany, Scaleway SAS in France and DataCrunch Oy (Verda) in Finland for inference, and requires adequacy or Standard Contractual Clauses for any sub-processor outside the EEA. The home page states that requests are not used for training. The privacy policy read covers the website and is available in German only, and no prompt-retention statement was found on the pages read.
Pay-as-you-go prices are per million tokens in euro, for example €0.20 input and €1 output for DeepSeek V4.1 Flash. The coding plans are Basic at €49, Pro at €149, Unlimited 3x at €499 and Unlimited 6x at €999 per month, offered only to businesses under § 14 BGB, prices plus VAT.
What Melious AI does well
- German GmbH with a commercial register number and named managing director
- Sub-processors named with country in the DPA
- Euro pay-as-you-go prices and published business plans
- OpenAI- and Anthropic-compatible endpoints for common coding tools
Where Melious AI falls short
- No statement on prompt retention was found on the pages read
- The privacy policy read covers the website and is in German only
- Plans are offered to businesses only
- A gateway across 11 providers means the data path varies by route
Standout feature. It is the only provider here whose DPA names the underlying hosting providers, IONOS and Scaleway among them, as inference sub-processors.