Koliber v1.2

Koliber v1.2 is the second full pretrained generation of the Koliber family, developed by OrisTeam and trained from scratch as a compact Polish decoder-only causal language model.

Koliber v1.2 keeps the core architectural direction of Koliber v1.0 while rebuilding the tokenizer, data pipeline, corpus mixture, document construction, and exposure recipe.

Koliber v1.2 was designed to test how far the existing Koliber architecture can be pushed through a better training stack before increasing model scale.

It is a base pretrained model, not an instruction-tuned assistant or chatbot. It is intended primarily for text completion and as a foundation for supervised fine-tuning, preference optimization, continued pretraining, research, and experimentation.

Why v1.2 was rebuilt

Koliber v1.2 went through two substantially different training attempts.

The first v1.2 candidate increased the model to approximately 175M parameters. It reached roughly 1.8B training tokens, but was stopped before full saturation. Although it could perform well on some categorical evaluations, its raw generation quality was not consistently better than Koliber v1.0 and remained behind stronger Polish baselines in blind generation tests.

More importantly, scaling the model did not remove the problems inherited from the v1.0 data pipeline. The larger candidate was still trained on a stack with weak control over final source exposure, imperfect document boundaries, and several corpus-quality issues.

For the final v1.2 run, OrisTeam therefore returned to the proven 126M-class Koliber core and focused on rebuilding the tokenizer, data pipeline, and training recipe first.

The larger-model direction is deferred until the data and training stack can demonstrate a reliable generative improvement.

Model overview

Property Koliber v1.2
Parameters ~126M
Decoder layers 12
Hidden size 768
Query heads / KV heads 12 / 2
Head dimension 64
FFN 3072
Context length 1536
Attention GQA
Positional encoding RoPE
Activation SwiGLU
Normalization RMSNorm
Bias No
LM head tied embeddings
Tokenizer Koliber v1.2 32K
Language primarily Polish
Stage base pretraining

Auxiliary objective

Koliber v1.0 used a small training-only future-state auxiliary objective that encouraged the model to predict a representation further ahead in the sequence. Koliber v1.2 does not use this auxiliary loss.

With v1.2, the training stream is built around explicit <s> DOCUMENT </s> boundaries and a stronger assumption that each packed segment should preserve the structure of an individual document. The auxiliary objective may still be useful for longer-range coherence within the 1536-token context, but v1.2 focuses first on correcting the data-mixture, document-boundary, tokenizer, and exposure issues observed in v1.0.

Selanka and web reconstruction

The web corpus was reprocessed using Selanka, an internal OrisTeam masked-language model used for document-quality and boundary decisions.

Its training data was prepared with several teacher models under a shared labeling policy, with the approximate contribution order:

DeepSeek v4.1 Flash > Gemma 4 26B A4B QAT > Bielik v3 11B Q8

This ordering describes dataset contribution, not a ranking of general model quality.

The Selanka classification dataset contained approximately 8,000 training examples. Around 20% of the training annotations were manually reviewed, with approximately 95% of the reviewed labels matching the intended policy.

The four operational decisions were:

Decision Meaning
Clean high-quality text requiring no structural correction
Keep usable text retained in the corpus
Split a document or quality boundary should be introduced
Drop material should not enter the training corpus

Selanka was fine-tuned using stepped class exposure:

4ร— Clean / 3ร— Keep / 2ร— Split / 1ร— Drop

The training set deliberately contained noisy and ambiguous material so that the model learned document-quality boundaries rather than only a clean-vs-noise distinction.

As an internal reference, Selanka was also evaluated against Polish RoBERTa-v2 Base. The models were approximately comparable in the local evaluation, while Selanka provided roughly 30% higher end-to-end throughput in the target processing stack.

Selanka was then used to re-evaluate and reconstruct the earlier web corpus, producing the web pools used by Horyzont v1.5.

Selecting the training corpus

Horyzont dataset v1.5 started from approximately 10.5B available tokens and selected roughly 4.2B tokens for the final training plan.

Selection used several signals together:

  • document deduplication,
  • language verification for the rebuilt web data, PPDC, and Biblioteka Nauki,
  • embedding-based similarity checks across corpora to reduce semantic overlap between sources,
  • source and document metadata,
  • lexical rarity,
  • perplexity-derived difficulty signals.

Perplexity-related metadata was produced with Bielik v3 11B Q8 and distilled into the lighter Selanka-based processing stack.

Difficulty was not defined by perplexity alone. Horyzont also estimated lexical rarity at the word level, while taking into account segmentation under both the Bielik and Koliber v1.2 tokenizers.

Conceptually, these signals form a small difficulty grid:

Lower perplexity Higher perplexity
Common vocabulary easier / regular harder structure or content
Rare vocabulary uncommon but predictable rare and difficult

The training stream starts with a stronger share of easier material, gradually introduces more rare and difficult text, and moves back toward more representative intermediate difficulty near the end instead of finishing exclusively on the hardest tail.

Quality and difficulty are separate signals. High perplexity or rare vocabulary does not by itself make a document better training data.

Data exposure: v1.0 vs v1.2

The largest practical change is how the model receives data.

Koliber v1.0

Koliber v1.0 selected the next source using relative sampling weights:

Source v1.0 sampling weight
Clean web 75%
Keep web 15%
Wikipedia 6%
Split web 2%
NKJP Balanced 1%
Judicial / legal 0.5%
OpenSubtitles 0.5%

These values controlled the probability of selecting a source for the next sequence. They did not guarantee the final number of tokens consumed from each corpus.

Sources could finish at different times, after which the remaining weights were renormalized. As a result, configured sampling probability, available corpus size, and actual training exposure were not the same thing.

Koliber v1.2

Horyzont v1.5 instead assigns an explicit token target to every training source:

Source Target exposure Share
Clean web v1.5 1.192B ~28.2%
Keep web v1.5 250M ~5.9%
Polish public-domain corpus 950M ~22.5%
Biblioteka Nauki 863M ~20.4%
Wikipedia 794M ~18.8%
Sejm 182M ~4.3%
Total 4.231B 100%

A source may be cycled for additional epochs when its planned exposure exceeds its unique physical size.

This separates corpus size from training exposure: large sources no longer dominate simply because more raw text is available, while smaller curated corpora can receive intentionally larger token budgets. Source priority may still affect when material appears during training, but not its final assigned exposure.

Training progress and token budget

The full v1.2 exposure recipe contains approximately 4.231B training tokens. The currently released checkpoint, update_00045529.pt, was trained for 2,797,301,760 tokens, corresponding to approximately 66.1% of the planned exposure recipe.

Reference point Training tokens Tokens / parameter (~126M) Share of v1.2 exposure plan Status
~20 tokens / parameter reference ~2.52B ~20.0ร— ~59.6% compute-optimal scaling reference
update_00045529.pt 2.797B ~22.2ร— ~66.1% current release
Full v1.2 exposure target 4.231B ~33.6ร— 100% planned training exposure

The often-cited ~20 training tokens per parameter figure is included only as a rough scaling-law reference, not as a hard threshold for when a model is "fully trained" or saturated. The released v1.2 checkpoint had already passed that reference before publication.

Training continued beyond this point because the v1.2 run was defined by its fixed source-exposure recipe rather than by the 20-token/parameter heuristic. Additional training can still change generative behavior, domain balance, and downstream usefulness even after that reference has been crossed.

Accordingly, 66.1% here means 66.1% of the planned v1.2 data-exposure schedule, not 66.1% of some assumed minimum amount of training required for the architecture.

Document packing and training stack changes

Koliber v1.0 primarily separated packed documents with an end-of-document token.

Koliber v1.2 explicitly reconstructs document boundaries as:

<s> DOCUMENT </s>

This change accompanies the Selanka-based document reconstruction and makes both the start and end of individual documents visible to the model.

The v1.2 training implementation also defaults to explicit K/V expansion for GQA before standard scaled-dot-product attention. On the tested NVIDIA GPU stack this was substantially faster than the native GQA path while preserving the same model parameterization.

Horyzont v1.5 โ†’ v2

Horyzont v1.5 is an extension of the existing OrisTeam data stack rather than its final form.

Its token mixture is the closest practical approximation of the intended training distribution that could be built from the available corpora, metadata, and filtering infrastructure without first materializing and maintaining a separate heavily weighted dataset on the order of tens of gigabytes.

Horyzont v2 is planned as a larger iteration of the system with dedicated tooling for inspecting and testing data recipes before committing to a full pretraining run.

The planned workflow includes a dashboard for comparing source distributions, exposure targets, curriculum choices, and the effective token stream produced by candidate recipes.

Generation

Koliber v1.2 is a base text-completion model. Natural text prefixes are generally more appropriate than instruction-style chat prompts.

The evaluated Koliber v1.2 checkpoint was update_00045529.pt, trained for 2,797,301,760 tokens.

Generation settings:

max_new_tokens = 96
temperature = 0.8
top_k = 40
top_p = 0.95
repetition_penalty = 1.15

Koliber v1.0 โ†’ v1.2

Koliber v1.2 largely removes the repetition collapse seen in v1.0. In the 96-prompt blind suite, a simple repetition check found at least three direct repetitions of the same token in 23/96 v1.0 generations, including 13/96 cases with at least ten repetitions. The same check found 0/96 such cases for v1.2.

A representative v1.0 failure looked like:

tu tu tu tu tu tu tu tu...

v1.2 more often produces complete sentences, preserves the prompt domain, and maintains readable Polish over longer continuations. Its main remaining weakness is semantic drift: the model may stay stylistically and topically plausible while losing the specific requirement or inventing dates, institutions, locations, quotations, and article-like details.

Koliber v1.0: language itself may collapse.
Koliber v1.2: language is much more stable, but semantics can still drift.

96-prompt blind comparison

The final blind generation test used 96 different prompts and compared:

  • Koliber v1.2 โ€” update_00045529.pt
  • Koliber v1.0 Base
  • GoLLeM v4 250M as a strong Polish reference model
Model 1st place 2nd place 3rd place Mean rank
GoLLeM v4 250M 76 / 96 (79.2%) 16 4 1.25
Koliber v1.2 14 / 96 (14.6%) 64 18 2.04
Koliber v1.0 6 / 96 (6.3%) 16 74 2.71

The result is consistent with a substantial improvement from v1.0 to v1.2. Koliber v1.2 most often occupies the middle position: it is much more stable and coherent than v1.0, but still clearly behind the strongest external reference in long-range semantic consistency.

Limitations and safety

Koliber v1.2 has not been supervised instruction-tuned and has not undergone preference optimization or chat alignment.

Compared with Koliber v1.0, raw generation is more stable and severe repetition failures are substantially reduced, but important limitations remain.

Raw generations may:

  • drift semantically while remaining locally fluent,
  • lose specific prompt constraints or gradually move away from the original topic,
  • hallucinate facts, citations, people, organizations, dates, locations, quotations, or statistics,
  • produce malformed, biased, offensive, unsafe, or otherwise undesirable text,
  • imitate patterns present in web, news, academic, legal, or other pretraining material,
  • behave unpredictably outside distributions represented during pretraining.

The released checkpoint should not be treated as the maximum capability of the architecture. It was released after approximately 2.797B training tokens (~22.2 tokens per parameter), while the full v1.2 exposure plan extends to approximately 4.231B tokens (~33.6 tokens per parameter). These figures describe training exposure and should not be interpreted as fixed saturation thresholds.

The model has also not undergone task-specific fine-tuning, so raw base-model generation should not be treated as the expected ceiling of downstream models built from the checkpoint.

Do not rely on raw generations for medical, legal, financial, safety-critical, or other high-stakes decisions.

Developers building downstream applications should perform their own evaluation, filtering, alignment, and domain-specific validation.

Intended use

Koliber v1.2 is intended for:

  • Polish text-completion experiments,
  • continued pretraining,
  • supervised fine-tuning,
  • preference optimization,
  • representation and likelihood-based evaluation,
  • research on compact Polish language models.

It is not presented as a safety-aligned assistant.

Citation

@misc{KoliberV12,
  author       = {Aleksander Ogrodzki},
  title        = {Koliber v1.2},
  year         = {2026},
  publisher    = {Hugging Face},
  url          = {https://huggingface.co/OrisTeam/Koliber-1.2},
  note         = {Model architecture, training data pipeline, and training pipeline developed by the author}
}

License

Apache-2.0.

Downloads last month
178
Safetensors
Model size
0.2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including OrisTeam/Koliber-1.2