Mijatovic-Adriaticum-1B

Mijatovic-Adriaticum-1B is an independent compact 1-billion-parameter causal language model developed by Mijatovic with a primary focus on Serbian, Croatian, Bosnian, and the wider South Slavic language space. Trained from scratch from random initialization, it is engineered for efficient text generation, instruction following, and conversational assistance.

Designed for everyone.


Model Details

  • Developer: Mijatovic
  • Model Name: Mijatovic-Adriaticum-1B
  • Hugging Face Repository: mijatovic/Adriaticum-1
  • API Model ID: adriaticum-1
  • Model Type: Custom causal language model (MijatovicForCausalLM)
  • Total Parameters: 1,025,541,888 (~1.03B / 1B class)
  • Context Length: 2,048 tokens (max_position_embeddings: 2048)
  • Vocabulary Size: 50,000 tokens
  • Precision: Float16 (torch.float16)
  • Weight Shards: 2 SafeTensors shards (~2.05 GB total)
  • Custom Modeling Code: configuration_mijatovic.py, modeling_mijatovic.py (trust_remote_code=True required)
  • License: Apache License 2.0
  • Official Website: https://mijatovic.io
  • Hugging Face Organization: https://huggingface.co/mijatovic

Languages

The primary public scope of Mijatovic-Adriaticum-1B is centered on:

  1. Serbian (both Latin and Cyrillic scripts)
  2. Croatian
  3. Bosnian

English-language text formed part of the foundational pretraining, instruction tuning, and translation data alignment, enabling functional cross-lingual understanding and translation capabilities. However, English is not marketed as the primary target language for this release. While the model has broad exposure to South Slavic linguistic patterns, performance and stylistic coverage are strongest in Serbian, Croatian, and Bosnian, and may vary across other regional variants.


Intended Use

Mijatovic-Adriaticum-1B is designed for:

  • Local and efficient conversational text generation.
  • Developer experimentation, prototyping, and lightweight application integration.
  • Academic and computational research into compact multilingual language models.
  • English–Croatian and English–Serbian translation and cross-lingual understanding tasks.

Out-of-Scope and Prohibited Uses

  • Not intended for safety-critical systems, autonomous decision-making, medical advice, legal counsel, or financial planning without qualified professional oversight.
  • Should not be deployed without human supervision in high-stakes environments.
  • Tool calling is not supported in this release.

Limitations

  • Compact Scale: As a ~1B-parameter model, Adriaticum may generate incorrect, incomplete, inconsistent, or fabricated information compared to larger frontier models.
  • Instruction Adherence: While trained for conversational interactions, instruction following may be imperfect, particularly with complex negative constraints or lengthy multi-step instructions.
  • Factuality & Verification: Output should not be relied upon as an authoritative source of facts. Users should independently verify outputs for important decisions.
  • Language & Domain Variation: Output quality may vary depending on prompt structure, linguistic dialect, domain terminology, and context length.
  • Context Window: The model operates within a 2,048-token context window; inputs exceeding this budget require external truncation.

Usage

Mijatovic-Adriaticum-1B includes a validated Hugging Face chat template embedded directly in the repository tokenizer configuration. The documented workflow below was tested with Transformers 5.13.0 and PyTorch.

Because Adriaticum utilizes custom modeling architecture, trust_remote_code=True is required when loading both the tokenizer and the model.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "mijatovic/Adriaticum-1"

tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
    torch_dtype="auto",
)

messages = [
    {"role": "user", "content": "Napiši kratko objašnjenje šta je jezički model."}
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
)

outputs = model.generate(
    **inputs,
    max_new_tokens=160,
    eos_token_id=1,
    pad_token_id=2,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)

Training

Mijatovic-Adriaticum-1B was trained from scratch through three distinct stages:

  1. Stage-1 Pretraining: Foundational causal language modeling from random initialization across large-scale web and textual corpora.
  2. Phase-A Language Adaptation & Replay: Focused language adaptation using regional parliamentary, legislative, and web text, combined with replay buffers.
  3. Phase-B Supervised Fine-Tuning (SFT): Instruction following, English–Croatian translation pairs, conversational alignment, and safety tuning.

Training Data

The model was trained on a curated mixture of public datasets, synthetic alignment data, and first-party alignment data:

  • FineWeb & FineWeb-Edu: Large-scale educational and web pretraining text (Hugging Face / Common Crawl).
  • CLASSLA-web 2.0: Web corpora for Bosnian, Croatian, and Serbian (Jožef Stefan Institute / CLARIN.SI).
  • ParlaMint (BA, HR, RS) v5.0: Parliamentary debate corpora for Bosnia and Herzegovina, Croatia, and Serbia (CLARIN ERIC).
  • DGT Translation Memory: Official English–Croatian parallel translation segments (European Commission, Directorate-General for Translation / JRC, Release 2021 / 2020 data update).
  • OpenAssistant OASST2: Filtered English instruction and dialogue pairs.
  • NVIDIA HelpSteer: Filtered English alignment data.
  • CohereLabs Aya Dataset: Filtered Serbian-language instruction subset.
  • Synthetic Alignment Data: Generated using DeepSeek Open Platform models deepseek-v4-flash and deepseek-v4-pro, then filtered, normalized, and deduplicated by Mijatovic before training.
  • Mijatovic Identity & Behavior v1: First-party alignment data developed by Mijatovic.

Mijatovic does not redistribute raw third-party training datasets in this repository, and Apache-2.0 does not relicense those datasets. For source-specific attributions, license terms, and legal notices, see THIRD_PARTY_NOTICES.md.

Privacy & User Data Assurance

Production Mijatovic user conversations were not used in training this release. All conversational samples derived exclusively from documented public datasets and curated first-party synthetic alignment pipelines.


Synthetic Data Disclosure

A subset of synthetic instruction/alignment data was generated using DeepSeek Open Platform models deepseek-v4-flash and deepseek-v4-pro, then filtered, normalized, and deduplicated before training. No affiliation with or endorsement by DeepSeek is implied.


Evaluation

Internal evaluation is ongoing. No public benchmark results are reported for this release.


License & Third-Party Notices

The Mijatovic-authored material distributed in this repository, including the model weights and custom modeling code, is released under the Apache License 2.0. See LICENSE and NOTICE.

Third-party training sources remain subject to their respective terms and notices. See THIRD_PARTY_NOTICES.md.


Links

Downloads last month
109
Safetensors
Model size
1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support