Genesis (genesis-4B-DEMO)

genesis_logo

Genesis is a a fine-tune variant of unsloth/Qwen3.5-4B, fine-tuned on a curated dataset of 1,171 Macedonian mathematical competition problem-solution pairs. This release marks the beginning of a bigger project, and a new series of language models, enhancing mathematical reasoning capabilities in under-resourced languages.

Model Details

  • Base Model: unsloth/Qwen3.5-4B
  • Adapter Type: LoRA (Rank r=16, Alpha α=32, Dropout 0.0)
  • Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Precision: bf16
  • Max Sequence Length: 1024
  • Training Epochs: 10 (peak performance observed at epoch 3)
  • Hardware: Single NVIDIA RTX 3060 12GB

Training Data

The model was fine-tuned on a manually verified dataset of Macedonian mathematical competition problems. The pipeline converts public PDFs into structured LaTeX, extracts problem-solution pairs, and applies rigorous manual verification. The final dataset contains 1,171 entries covering algebra, number theory, combinatorics, and arithmetic reasoning. Geometry problems requiring figure reconstruction were excluded due to pipeline limitations.

Dataset Availability: The training dataset is not publicly released. Only the model weights and training/evaluation code are provided.

Evaluation

Benchmark Base Qwen3.5-4B Genesis (Epoch 3) Genesis (Epoch 10)
GSM8K_mk (748 problems) 9.90% 57.75% 44.92%
macedonian-llm-eval (avg) 0.47 0.47 0.47
  • GSM8K_mk: Macedonian translation of GSM8K. Pass@1 with greedy decoding and regex-based answer extraction.
  • macedonian-llm-eval: Seven standard benchmarks (arc_challenge, arc_easy, boolq, hellaswag, openbookqa, piqa, winogrande). Zero-shot, greedy decoding.
  • Key Finding: Performance peaks at epoch 3 and degrades monotonically with extended training, despite continued reduction in training loss. General Macedonian capabilities are preserved.

genesis_gsm8k_accuracy Pass@1 accuracy on the 748-problem Macedonian GSM8K test set across Genesis checkpoints and baseline models. Performance peaks at epoch 3 and declines monotonically.

genesis_macedonian_eval Performance comparison across the seven macedonian-llm-eval benchmarks. Genesis (E3) maintains parity with the base model, confirming preservation of general Macedonian capabilities.

Usage

Prerequisites

pip install unsloth transformers torch datasets peft accelerate

Loading the Adapter

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "a-nikoloski/genesis-4b-demo"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "Пресметај го збирот на броевите од 1 до 10."}]
inputs = tok.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True,
    return_tensors="pt", return_dict=True,
).to(model.device)

with torch.no_grad():
    out = model.generate(
        **inputs, max_new_tokens=512, do_sample=False,
        repetition_penalty=1.1, pad_token_id=tok.pad_token_id or tok.eos_token_id,
    )

print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Citation

If you use this model or the associated training/evaluation code, please cite:

@misc{genesis2026,
  author = {Aleksandar Nikoloski},
  title = {Genesis: Towards Macedonian Mathematical Language Models},
  year = {2026},
  url = {https://huggingface.co/epsill0n/genesis-4B-DEMO}
}

License

This model is released under the Apache License 2.0. See LICENSE for details.

Downloads last month
732
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for epsill0n/genesis-4B-DEMO

Finetuned
Qwen/Qwen3.5-4B
Adapter
(173)
this model

Evaluation results