Hacker News

yu3zhou4
Kolibri – Tech Report [pdf] aleph-alpha.com

senko3 hours ago

Since it's a PDF, here's an excerpt from the summary:

> This report introduces Kolibri, an English–German Mixture-of-Experts transformer with 78.1B total parameters and 3.46B active parameters per token, released as open weights under the Apache 2.0 licence.

> We built Kolibri for sovereign and specialised deployment, with a particular focus on German and on regulated domains such as public administration, industry, and aerospace. To support these settings, we control the full model-development process, including data, architecture, training infrastructure, post-training, and evaluation.

Dupe of https://news.ycombinator.com/item?id=49942706

yu3zhou4op2 hours ago

It’s not a dupe, it’s a full training and infra report

peterBlue75an hour ago

Related thread on the tooling:

Model Training as Code - https://news.ycombinator.com/item?id=48673450 - June 2026 (24 comments)

satvikpendeman hour ago

And related thread on the actual model post: https://news.ycombinator.com/item?id=49942706

derin-picment32 minutes ago

Skimmed the first sections — the most interesting part to me isn't just the 78.1B total / 3.46B active MoE numbers, but the data story: 24T tokens with >20% German, including 2T+ German tokens curated/generated themselves.

That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.

Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.

hn-front (c) 2024 voximity
source