The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
andaitoday at 5:34 PM
>We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context.
The thing to note here, besides the transparency and the fact that itās actually a good model that also works well on coding and agentic tasks, is that itās the first release by a team formed less than a year ago, with a strong focus on iteration velocity. Thereās more to come.
disclaimer: Iām part of the training team, happy to answer any questions
tomCombtoday at 4:05 PM
For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
amoshebbtoday at 2:50 PM
Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
niemandhiertoday at 2:07 PM
I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a ātrust adapterā.
If the second model is cheap and fast enough, there is a business model.
You donāt even need to audit all the intermediate steps, just tool calls and end results.
spijdartoday at 2:00 PM
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.
I get that doesn't invalidate the real "point" of the model, but...
martianvoidtoday at 12:56 PM
I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if itās able to catch the correct approach
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.
This thing is worse than a Qwen3.8 27B.
gpugregtoday at 8:57 PM
I was wondering whether the model could help me with German bureaucracy. Unfortunately, the answer is "no".
More specifically, I asked the model about what I should put on my contact page, which can cost you in the order of 500 ⬠in Germany if you don't write the right magic words.
Kolibri incorrectly referenced the "Telemediengesetz" ("telecommunication act"), which has been superseded by the "Digitale-Dienste-Gesetz" (DDG, "digital services act") since 2024. The model knows about the DDG, but does not reference it unless specifically instructed to do so.
If anyone of the developers reads this, you can fix this by introducing a recency bias during training. You can even control it by conditioning the model on a date provided with the system prompt or first prompt, so you can travel in time.
mark_l_watsontoday at 3:57 PM
Looks interesting. I just went to download from HF, but they only have fp16 which won't fit on my Mac.
Good to see Europe adding toe what Mistral is doing. +100
vzalivatoday at 3:28 PM
"Languages: German and English" ā this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace
Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...
james45today at 3:11 PM
the sovereignty topic needs more attention in general so great to see. self-hosting the model is one piece of sovereignty, but how do we handle the rest of the agent stack - embeddings, retrieval, memory, etc. Has anyone put together a practical agent stack that's 100% sovereign, where they control it all?
lmf4loltoday at 9:25 PM
Wow this is so cool. Glad that Aleph Alpha does that after Mistral threw the towel in the ring (and disappointed the european AI crowd massively!!!!). After AA got sold to the Canadians, I thought its over but this is a really cool comeback and the depth of the tech report shows that they a serious about openness.
I hope I can use their model soon in my product. Would be awesome to have a European model to offer!!!!
I really wonder how it compares to deepseek v4.1 flash
cheesecakegoodtoday at 3:56 PM
For those sick of āPareto frontierā talk, just shorthand it as āitās the best at some very particular thingā. Obviously that one thing/tradeoff itās good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that theyāve chased some tiny edge into the ground.
Iāll be curious to see which it becomes in the next year - nba āvery narrow recordā, or a sign you can hang with the big boys.
UncleOxidanttoday at 7:31 PM
78B MoE with A3.6B is a very nice size.
x1watttoday at 1:32 PM
Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.
dosingatoday at 3:25 PM
> The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
petesergeanttoday at 1:55 PM
I wish nothing but luck for an EU model, but:
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
wg0today at 6:38 PM
Anyone thinking this won't improve or is behind etc is blatantly wrong. It'll catchup within a year. Like that unknown wise and visionary man inside Google once said about their competitors: We have no moat neither does anyone else."
Congrats to the team.
toshtoday at 2:06 PM
i wonder if the custom tokenizer is better in practice, the examples look interesting though
Jeeetendratoday at 3:57 PM
3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
pythonic_helltoday at 1:06 PM
The benchmarks are impressive given the problem space they are working in.
d2kxtoday at 1:19 PM
German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
CorezIoOfficialtoday at 2:09 PM
Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
erelongtoday at 4:42 PM
is this like an unfortunate name clash with KolibriOS (kind of like how Google Gemini was a clash with the Gemini protocol project)?
obliotoday at 2:39 PM
I wonder if we can start having LLM distros: community led distributed training runs with periodic releases, open weights, FOSS code, the whole shebang. Maybe the public training sets can reach a level where an LLM trained on them can be good enough for most things, such as web search and aggregation, coding, etc.
I wonder how far we are from this. How far are we from LLM's Debian moment?
JaggerJotoday at 3:51 PM
Is this a truely open source model or also open weights?
woadwarrior01today at 1:31 PM
> A bigger dense model beats it. Qwen3.8 27B ...
How is a 27B dense model bigger than a 78B MoE?
cbarricktoday at 1:05 PM
I got distracted by that scroll-wheel UI component on the page. Neat!
orifitotoday at 1:53 PM
At least Germany is moving smarter than UK government...
veryfancytoday at 1:55 PM
Nice to see public goods in this space.
ThouYStoday at 2:44 PM
calling qwen 27B a bigger model.. I don't know man. My vram says otherwise.
deletedtoday at 1:32 PM
deletedtoday at 6:58 PM
cyanydeeztoday at 12:57 PM
Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware.
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
mistyvalestoday at 1:34 PM
The Sega 32X game??
Larrikintoday at 4:39 PM
The name really evokes strong Kotlin library naming vibes.
12949468today at 4:02 PM
It still steals my IP without attribution. Now we have state sanctioned sovereign theft instead of foreign theft.
9devtoday at 1:06 PM
Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.
sajithdilshantoday at 12:15 PM
> It knows less from memory, Multi-turn tool calling is weaker, Itās not the best coding agent
Then what does it good at? Sending faxes?
hypfertoday at 2:00 PM
The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd.
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
gilfoyle_7today at 3:22 PM
does it have GDPR compliance?
shevy-javatoday at 2:59 PM
Is Germany sovereign? It outsourced its defence onto the USA. Recently Trump wanted more diesel; Germany insta-submitted, also because oddly enough Macron submitted before Germany (Macron is suspicious). Before that, Leyen committed to insta-submission with a deal that made europeans poorer (and perhaps Leyen benefits from that). Canada shows the way. Many of the smaller countries in the EU too, such as Netherlands, Denmark, Finland, to some extent Sweden as well. Every time I read "sovereign" here I have to object. Nothing is sovereign here. The whole hardware is definitely not sovereign. Perhaps some of the software is, but that's about it. Plus, who gets all the data? The big US mega-corporations sniff non-stop. Remember how Facebook sniffed Libgen and Anna's Archive dry etc..., then suddenly libgen went down. The US corporations act as huge global leeches on every step of the stair. And lobbyists benefit from this too.
Lucasoatotoday at 1:28 PM
> 4. It thinks in German
This means that itās always on time, it uses acronyms for everything and when thereās a decision to be made, it sets up a committee.
tfburnstoday at 5:37 PM
[flagged]
derin-picmenttoday at 8:47 PM
Skimmed the first sections ā the most interesting part to me isn't just the 78.1B total / 3.46B active MoE numbers, but the data story: 24T tokens with >20% German, including 2T+ German tokens curated/generated themselves.
That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest ā serving cost matters a lot for regulated on-prem use.
Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.