Instructions to use OrisTeam/Koliber-1.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OrisTeam/Koliber-1.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OrisTeam/Koliber-1.2", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("OrisTeam/Koliber-1.2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OrisTeam/Koliber-1.2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OrisTeam/Koliber-1.2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrisTeam/Koliber-1.2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/OrisTeam/Koliber-1.2
- SGLang
How to use OrisTeam/Koliber-1.2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OrisTeam/Koliber-1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrisTeam/Koliber-1.2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OrisTeam/Koliber-1.2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrisTeam/Koliber-1.2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use OrisTeam/Koliber-1.2 with Docker Model Runner:
docker model run hf.co/OrisTeam/Koliber-1.2
Koliber v1.2
Koliber v1.2 is the second full pretrained generation of the Koliber family, developed by OrisTeam and trained from scratch as a compact Polish decoder-only causal language model.
Koliber v1.2 keeps the core architectural direction of Koliber v1.0 while rebuilding the tokenizer, data pipeline, corpus mixture, document construction, and exposure recipe.
Koliber v1.2 was designed to test how far the existing Koliber architecture can be pushed through a better training stack before increasing model scale.
It is a base pretrained model, not an instruction-tuned assistant or chatbot. It is intended primarily for text completion and as a foundation for supervised fine-tuning, preference optimization, continued pretraining, research, and experimentation.
Why v1.2 was rebuilt
Koliber v1.2 went through two substantially different training attempts.
The first v1.2 candidate increased the model to approximately 175M parameters. It reached roughly 1.8B training tokens, but was stopped before full saturation. Although it could perform well on some categorical evaluations, its raw generation quality was not consistently better than Koliber v1.0 and remained behind stronger Polish baselines in blind generation tests.
More importantly, scaling the model did not remove the problems inherited from the v1.0 data pipeline. The larger candidate was still trained on a stack with weak control over final source exposure, imperfect document boundaries, and several corpus-quality issues.
For the final v1.2 run, OrisTeam therefore returned to the proven 126M-class Koliber core and focused on rebuilding the tokenizer, data pipeline, and training recipe first.
The larger-model direction is deferred until the data and training stack can demonstrate a reliable generative improvement.
Model overview
| Property | Koliber v1.2 |
|---|---|
| Parameters | ~126M |
| Decoder layers | 12 |
| Hidden size | 768 |
| Query heads / KV heads | 12 / 2 |
| Head dimension | 64 |
| FFN | 3072 |
| Context length | 1536 |
| Attention | GQA |
| Positional encoding | RoPE |
| Activation | SwiGLU |
| Normalization | RMSNorm |
| Bias | No |
| LM head | tied embeddings |
| Tokenizer | Koliber v1.2 32K |
| Language | primarily Polish |
| Stage | base pretraining |
Auxiliary objective
Koliber v1.0 used a small training-only future-state auxiliary objective that encouraged the model to predict a representation further ahead in the sequence. Koliber v1.2 does not use this auxiliary loss.
With v1.2, the training stream is built around explicit <s> DOCUMENT </s> boundaries and a stronger assumption that each packed segment should preserve the structure of an individual document. The auxiliary objective may still be useful for longer-range coherence within the 1536-token context, but v1.2 focuses first on correcting the data-mixture, document-boundary, tokenizer, and exposure issues observed in v1.0.
Selanka and web reconstruction
The web corpus was reprocessed using Selanka, an internal OrisTeam masked-language model used for document-quality and boundary decisions.
Its training data was prepared with several teacher models under a shared labeling policy, with the approximate contribution order:
DeepSeek v4.1 Flash > Gemma 4 26B A4B QAT > Bielik v3 11B Q8
This ordering describes dataset contribution, not a ranking of general model quality.
The Selanka classification dataset contained approximately 8,000 training examples. Around 20% of the training annotations were manually reviewed, with approximately 95% of the reviewed labels matching the intended policy.
The four operational decisions were:
| Decision | Meaning |
|---|---|
| Clean | high-quality text requiring no structural correction |
| Keep | usable text retained in the corpus |
| Split | a document or quality boundary should be introduced |
| Drop | material should not enter the training corpus |
Selanka was fine-tuned using stepped class exposure:
4ร Clean / 3ร Keep / 2ร Split / 1ร Drop
The training set deliberately contained noisy and ambiguous material so that the model learned document-quality boundaries rather than only a clean-vs-noise distinction.
As an internal reference, Selanka was also evaluated against Polish RoBERTa-v2 Base. The models were approximately comparable in the local evaluation, while Selanka provided roughly 30% higher end-to-end throughput in the target processing stack.
Selanka was then used to re-evaluate and reconstruct the earlier web corpus, producing the web pools used by Horyzont v1.5.
Selecting the training corpus
Horyzont dataset v1.5 started from approximately 10.5B available tokens and selected roughly 4.2B tokens for the final training plan.
Selection used several signals together:
- document deduplication,
- language verification for the rebuilt web data, PPDC, and Biblioteka Nauki,
- embedding-based similarity checks across corpora to reduce semantic overlap between sources,
- source and document metadata,
- lexical rarity,
- perplexity-derived difficulty signals.
Perplexity-related metadata was produced with Bielik v3 11B Q8 and distilled into the lighter Selanka-based processing stack.
Difficulty was not defined by perplexity alone. Horyzont also estimated lexical rarity at the word level, while taking into account segmentation under both the Bielik and Koliber v1.2 tokenizers.
Conceptually, these signals form a small difficulty grid:
| Lower perplexity | Higher perplexity | |
|---|---|---|
| Common vocabulary | easier / regular | harder structure or content |
| Rare vocabulary | uncommon but predictable | rare and difficult |
The training stream starts with a stronger share of easier material, gradually introduces more rare and difficult text, and moves back toward more representative intermediate difficulty near the end instead of finishing exclusively on the hardest tail.
Quality and difficulty are separate signals. High perplexity or rare vocabulary does not by itself make a document better training data.
Data exposure: v1.0 vs v1.2
The largest practical change is how the model receives data.
Koliber v1.0
Koliber v1.0 selected the next source using relative sampling weights:
| Source | v1.0 sampling weight |
|---|---|
| Clean web | 75% |
| Keep web | 15% |
| Wikipedia | 6% |
| Split web | 2% |
| NKJP Balanced | 1% |
| Judicial / legal | 0.5% |
| OpenSubtitles | 0.5% |
These values controlled the probability of selecting a source for the next sequence. They did not guarantee the final number of tokens consumed from each corpus.
Sources could finish at different times, after which the remaining weights were renormalized. As a result, configured sampling probability, available corpus size, and actual training exposure were not the same thing.
Koliber v1.2
Horyzont v1.5 instead assigns an explicit token target to every training source:
| Source | Target exposure | Share |
|---|---|---|
| Clean web v1.5 | 1.192B | ~28.2% |
| Keep web v1.5 | 250M | ~5.9% |
| Polish public-domain corpus | 950M | ~22.5% |
| Biblioteka Nauki | 863M | ~20.4% |
| Wikipedia | 794M | ~18.8% |
| Sejm | 182M | ~4.3% |
| Total | 4.231B | 100% |
A source may be cycled for additional epochs when its planned exposure exceeds its unique physical size.
This separates corpus size from training exposure: large sources no longer dominate simply because more raw text is available, while smaller curated corpora can receive intentionally larger token budgets. Source priority may still affect when material appears during training, but not its final assigned exposure.
Training progress and token budget
The full v1.2 exposure recipe contains approximately 4.231B training tokens. The currently released checkpoint, update_00045529.pt, was trained for 2,797,301,760 tokens, corresponding to approximately 66.1% of the planned exposure recipe.
| Reference point | Training tokens | Tokens / parameter (~126M) | Share of v1.2 exposure plan | Status |
|---|---|---|---|---|
| ~20 tokens / parameter reference | ~2.52B | ~20.0ร | ~59.6% | compute-optimal scaling reference |
update_00045529.pt |
2.797B | ~22.2ร | ~66.1% | current release |
| Full v1.2 exposure target | 4.231B | ~33.6ร | 100% | planned training exposure |
The often-cited ~20 training tokens per parameter figure is included only as a rough scaling-law reference, not as a hard threshold for when a model is "fully trained" or saturated. The released v1.2 checkpoint had already passed that reference before publication.
Training continued beyond this point because the v1.2 run was defined by its fixed source-exposure recipe rather than by the 20-token/parameter heuristic. Additional training can still change generative behavior, domain balance, and downstream usefulness even after that reference has been crossed.
Accordingly, 66.1% here means 66.1% of the planned v1.2 data-exposure schedule, not 66.1% of some assumed minimum amount of training required for the architecture.
Document packing and training stack changes
Koliber v1.0 primarily separated packed documents with an end-of-document token.
Koliber v1.2 explicitly reconstructs document boundaries as:
<s> DOCUMENT </s>
This change accompanies the Selanka-based document reconstruction and makes both the start and end of individual documents visible to the model.
The v1.2 training implementation also defaults to explicit K/V expansion for GQA before standard scaled-dot-product attention. On the tested NVIDIA GPU stack this was substantially faster than the native GQA path while preserving the same model parameterization.
Horyzont v1.5 โ v2
Horyzont v1.5 is an extension of the existing OrisTeam data stack rather than its final form.
Its token mixture is the closest practical approximation of the intended training distribution that could be built from the available corpora, metadata, and filtering infrastructure without first materializing and maintaining a separate heavily weighted dataset on the order of tens of gigabytes.
Horyzont v2 is planned as a larger iteration of the system with dedicated tooling for inspecting and testing data recipes before committing to a full pretraining run.
The planned workflow includes a dashboard for comparing source distributions, exposure targets, curriculum choices, and the effective token stream produced by candidate recipes.
Generation
Koliber v1.2 is a base text-completion model. Natural text prefixes are generally more appropriate than instruction-style chat prompts.
The evaluated Koliber v1.2 checkpoint was update_00045529.pt, trained for 2,797,301,760 tokens.
Generation settings:
max_new_tokens = 96
temperature = 0.8
top_k = 40
top_p = 0.95
repetition_penalty = 1.15
Koliber v1.0 โ v1.2
Koliber v1.2 largely removes the repetition collapse seen in v1.0. In the 96-prompt blind suite, a simple repetition check found at least three direct repetitions of the same token in 23/96 v1.0 generations, including 13/96 cases with at least ten repetitions. The same check found 0/96 such cases for v1.2.
A representative v1.0 failure looked like:
tu tu tu tu tu tu tu tu...
v1.2 more often produces complete sentences, preserves the prompt domain, and maintains readable Polish over longer continuations. Its main remaining weakness is semantic drift: the model may stay stylistically and topically plausible while losing the specific requirement or inventing dates, institutions, locations, quotations, and article-like details.
Koliber v1.0: language itself may collapse.
Koliber v1.2: language is much more stable, but semantics can still drift.
96-prompt blind comparison
The final blind generation test used 96 different prompts and compared:
- Koliber v1.2 โ update_00045529.pt
- Koliber v1.0 Base
- GoLLeM v4 250M as a strong Polish reference model
| Model | 1st place | 2nd place | 3rd place | Mean rank |
|---|---|---|---|---|
| GoLLeM v4 250M | 76 / 96 (79.2%) | 16 | 4 | 1.25 |
| Koliber v1.2 | 14 / 96 (14.6%) | 64 | 18 | 2.04 |
| Koliber v1.0 | 6 / 96 (6.3%) | 16 | 74 | 2.71 |
The result is consistent with a substantial improvement from v1.0 to v1.2. Koliber v1.2 most often occupies the middle position: it is much more stable and coherent than v1.0, but still clearly behind the strongest external reference in long-range semantic consistency.
Limitations and safety
Koliber v1.2 has not been supervised instruction-tuned and has not undergone preference optimization or chat alignment.
Compared with Koliber v1.0, raw generation is more stable and severe repetition failures are substantially reduced, but important limitations remain.
Raw generations may:
- drift semantically while remaining locally fluent,
- lose specific prompt constraints or gradually move away from the original topic,
- hallucinate facts, citations, people, organizations, dates, locations, quotations, or statistics,
- produce malformed, biased, offensive, unsafe, or otherwise undesirable text,
- imitate patterns present in web, news, academic, legal, or other pretraining material,
- behave unpredictably outside distributions represented during pretraining.
The released checkpoint should not be treated as the maximum capability of the architecture. It was released after approximately 2.797B training tokens (~22.2 tokens per parameter), while the full v1.2 exposure plan extends to approximately 4.231B tokens (~33.6 tokens per parameter). These figures describe training exposure and should not be interpreted as fixed saturation thresholds.
The model has also not undergone task-specific fine-tuning, so raw base-model generation should not be treated as the expected ceiling of downstream models built from the checkpoint.
Do not rely on raw generations for medical, legal, financial, safety-critical, or other high-stakes decisions.
Developers building downstream applications should perform their own evaluation, filtering, alignment, and domain-specific validation.
Intended use
Koliber v1.2 is intended for:
- Polish text-completion experiments,
- continued pretraining,
- supervised fine-tuning,
- preference optimization,
- representation and likelihood-based evaluation,
- research on compact Polish language models.
It is not presented as a safety-aligned assistant.
Citation
@misc{KoliberV12,
author = {Aleksander Ogrodzki},
title = {Koliber v1.2},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/OrisTeam/Koliber-1.2},
note = {Model architecture, training data pipeline, and training pipeline developed by the author}
}
License
Apache-2.0.
- Downloads last month
- 178