Pleias
Paris lab building small open LLMs and the largest rights-cleared training dataset in Europe
Quick Overview
| Company | Pleias SAS |
|---|---|
| Category | LLMs & AI |
| Headquarters | Paris, France |
| Founded | 2023 |
| EU Presence | EU (France) |
| Open Source | Apache 2.0 (Pleias small language models) |
| Pricing | Common Corpus dataset and Pleias small language models (Pleias-RAG-1B, Pleias-Pico and others) free and open on Hugging Face under Apache 2.0; Synth and Stratum on-premise deployments priced through a demo call, no published rate card. |
| Free Option | Free open models and Common Corpus dataset on Hugging Face; book a demo for Synth/Stratum |
| Replaces | Proprietary LLM training-data vendors, closed synthetic-data tools |
Detailed Review
Pleias occupies a different corner of the European AI map: instead of chasing frontier scale, it builds the data layer underneath — and increasingly, the small models trained on it.
Common Corpus, its flagship dataset, is described as the largest open, rights-cleared collection for LLM pre-training, built from government records, legal archives and scientific literature with clear provenance rather than scraped copyrighted text. The work won an ICLR Oral, and partners include Nvidia, Mozilla, the Wikimedia Foundation and the AI Alliance.
On top of that data, Pleias trains and publishes small language models — including Pleias-RAG-1B and the Pleias-Pico family — openly on Hugging Face under Apache 2.0, small enough to run offline on hardware costing under €100.
Sillon, a 600-million-parameter model built for Paris transport operator RATP, is a concrete production example of a small, specialised model reported to outperform general models many times its size on its specific task. Synth and Stratum extend the same approach into synthetic training data and document processing, both deployable fully on-premise.
The trade-off is scale and polish. Pleias is not building a ChatGPT competitor — there is no general-purpose flagship chat product, no public pricing page, and engaging Synth or Stratum means booking a demo rather than reading a price list. Founded in 2023 as a small Paris SAS, it is also young and thinly staffed next to Mistral. What it offers instead is a genuinely open, auditable data supply chain and small models with a traceable training history.
What Pleias does well
- Common Corpus: largest open, rights-cleared LLM training dataset
- Small open-weight models on Hugging Face under Apache 2.0
- Production track record (RATP's Sillon model)
- On-premise deployment for Synth and Stratum
- Backed by Nvidia, Mozilla and the Wikimedia Foundation as partners
Where Pleias falls short
- No general-purpose flagship chat product
- No public pricing for Synth or Stratum
- Small, young company founded in 2023
- Best fit is data and training pipelines, not an end-user assistant
Standout feature. Training data with known, rights-cleared provenance — the opposite of "trust us" that most LLM pre-training asks for.
Pros and Cons
Pros
- Common Corpus: largest open, rights-cleared LLM training dataset
- Small open-weight models on Hugging Face under Apache 2.0
- Production track record (RATP's Sillon model)
- On-premise deployment for Synth and Stratum
- Backed by Nvidia, Mozilla and the Wikimedia Foundation as partners
Cons
- No general-purpose flagship chat product
- No public pricing for Synth or Stratum
- Small, young company founded in 2023
- Best fit is data and training pipelines, not an end-user assistant
Alternatives to Pleias
Frequently Asked Questions
What is Pleias?
Pleias occupies a different corner of the European AI map: instead of chasing frontier scale, it builds the data layer underneath — and increasingly, the small models trained on it.
Common Corpus, its flagship dataset, is described as the largest open, rights-cleared collection for LLM pre-training, built from government records, legal archives and scientific literature with clear provenance rather than scraped copyrighted text. The work won an ICLR Oral, and partners include Nvidia, Mozilla, the Wikimedia Foundation and the AI Alliance.
Where is Pleias based?
Pleias operates from Paris, France, which places it under EU (France).
What does Pleias cost?
Common Corpus dataset and Pleias small language models (Pleias-RAG-1B, Pleias-Pico and others) free and open on Hugging Face under Apache 2.0; Synth and Stratum on-premise deployments priced through a demo call, no published rate card.. Free open models and Common Corpus dataset on Hugging Face; book a demo for Synth/Stratum.
Who is Pleias best for?
Teams needing open, rights-cleared training data and small models. Training data with known, rights-cleared provenance — the opposite of "trust us" that most LLM pre-training asks for.
What are the drawbacks of Pleias?
No general-purpose flagship chat product. No public pricing for Synth or Stratum. Small, young company founded in 2023. Best fit is data and training pipelines, not an end-user assistant.