Heabsy
Slovak inference API for open models, on the company's own EEA hardware with zero data retention
Quick Overview
| Company | FEYA, s.r.o. |
|---|---|
| Category | LLMs & AI |
| Headquarters | Bratislava-Rača, Slovakia |
| Founded | Not stated by the vendor |
| EU Presence | EU (Slovakia) |
| Data Location | EEA only for the flagship model (own hardware); routed catalog models may process outside the EEA via third-party providers |
| Open Source | Not stated; the platform is closed-source, the models it serves are open-weight |
| Compliance | GDPR-aligned DPA (Art. 28), zero data retention on the flagship model, no US parent company |
| Pricing | Pay-as-you-go; flagship model $0.04/M input, $0.04/M cached input, $0.30/M output; no minimum spend |
| Free Option | No free tier; no credit card required to get a spending-limited key |
| Replaces | OpenAI API, Anthropic API, OpenRouter |
Detailed Review
Heabsy sells access to AI models rather than a model of its own: an OpenAI- and Anthropic-compatible API, built so a team already using the OpenAI SDK or running Claude Code, Cursor or Cline can point their client at a new base URL and keep everything else unchanged. The company, FEYA, s.r.o., is registered in Bratislava, Slovakia, with no US parent, which keeps the operator itself outside the reach of the US CLOUD Act.
The distinction that matters most on this page is between the flagship model and the rest of the catalog.
Heabsy's own model, Qwen3.8 27B, runs on hardware the company owns and operates inside the European Economic Area, and its Data Processing Agreement is specific about what that buys: prompts and completions are never written to logs, nothing is used to train any model, and no request content leaves the EEA.
Everything else in the catalog — models from Z.AI, OpenAI's open weights, Mistral, DeepSeek, Google and Meta — is fulfilled through global providers under Heabsy's own European contract, on one invoice and one DPA, but the compute for those may run outside the EEA under that provider's own policy.
On performance, Heabsy publishes numbers rather than asking to be taken on faith: roughly 176 tokens per second on the flagship model, about double the fastest public provider of the same model on the OpenRouter showcase, with a time to first token of 0.3–0.9 seconds.
The gain comes from a trained speculative-decoding draft running on dedicated hardware, not from quietly serving a smaller model.
A prefix cache keeps a conversation pinned to the machine that already holds its context, which the company says covers 91–97% of tokens on agentic workloads and cuts a repeated 34,000-token prompt from 42 seconds to 0.48 — a difference an agent that iterates ten times on the same context will feel directly.
What Heabsy does well
- Slovak company with no US parent, and the flagship model runs on EEA-owned hardware with a documented zero-retention policy
- OpenAI- and Anthropic-compatible: point an existing SDK or agent tool at a new base URL, nothing else changes
- Published, verifiable performance numbers (throughput, time to first token, cache hit rate) rather than marketing claims
- Pay-as-you-go with no minimum spend and no subscription, and per-key budgets with early rate-limit signals instead of a silent queue
Where Heabsy falls short
- Only one model (Qwen3.8 27B) carries the full EEA-only, zero-retention guarantee; the rest of the catalog depends on third-party providers whose compute may sit outside the EEA
- Young platform with no published founding date or team page
- No free tier, and the flagship model's 262,144-token context is smaller than some routed alternatives in the catalog
Standout feature. The prefix cache: on agentic coding workloads where most tokens are rereads of context the model has already seen, Heabsy reports 91–97% of tokens served from cache rather than recomputed, and bills cached input at a fraction of the fresh-token rate.
Pros and Cons
Pros
- Slovak company with no US parent, and the flagship model runs on EEA-owned hardware with a documented zero-retention policy
- OpenAI- and Anthropic-compatible: point an existing SDK or agent tool at a new base URL, nothing else changes
- Published, verifiable performance numbers (throughput, time to first token, cache hit rate) rather than marketing claims
- Pay-as-you-go with no minimum spend and no subscription, and per-key budgets with early rate-limit signals instead of a silent queue
Cons
- Only one model (Qwen3.8 27B) carries the full EEA-only, zero-retention guarantee; the rest of the catalog depends on third-party providers whose compute may sit outside the EEA
- Young platform with no published founding date or team page
- No free tier, and the flagship model's 262,144-token context is smaller than some routed alternatives in the catalog
Alternatives to Heabsy
Frequently Asked Questions
What is Heabsy?
Heabsy is a Slovak AI inference platform: an OpenAI- and Anthropic-compatible API serving open models. Its flagship model runs on Heabsy's own GPUs inside the EEA with zero data retention - prompts and completions are never logged. A routed catalog of other open models is fulfilled through global providers under a European contract, with compute that may run outside the EEA.
Where is Heabsy based?
Heabsy is operated by FEYA, s.r.o., registered at Pekná cesta 19, 831 52 Bratislava-Rača, Slovak Republic (IČO 47 887 541), which places it under EU (Slovakia) law with no US parent company. Its Data Processing Agreement states that inference on Heabsy-served models is provisioned entirely within the EEA.
What does Heabsy cost?
Pay-as-you-go, no minimum spend and no subscription. The flagship model (Qwen3.8 27B, run on Heabsy's own EEA hardware) costs $0.04 per million input tokens, $0.04 per million cached input tokens, and $0.30 per million output tokens. Routed catalog models are priced separately and published live, starting from $0.04 per million input tokens.
Who is Heabsy best for?
Development teams and agent builders - using tools like Claude Code, Cursor or Cline - who want an OpenAI- or Anthropic-compatible endpoint on European infrastructure with zero prompt retention, without operating their own GPU inference stack.
What are the drawbacks of Heabsy?
Only one model (Qwen3.8 27B) runs on Heabsy's own EEA hardware with the zero-retention guarantee; the rest of the catalog is fulfilled through third-party providers whose compute may sit outside the EEA. Young platform with no published founding date or team page, and the flagship model's context is capped at 262,144 tokens versus over a million on some routed alternatives.
Is Heabsy a good alternative to OpenAI or Anthropic's own APIs?
Heabsy is built as a European alternative for API access to AI models: a Slovak, OpenAI- and Anthropic-compatible inference API running open models on EEA-owned hardware with zero data retention. It will not be a like-for-like feature match in every respect, so check the review above for where it genuinely differs before switching.
How does Heabsy compare to other LLMs & AI tools?
Heabsy is one of several European LLMs & AI tools covered on this site. See the full LLMs & AI category page for how it ranks against its closest European competitors.