Tyrian 500M

A 511M parameter decoder-only language model built from scratch in PyTorch, then supervised fine-tuned for chat (ChatML) and preference-tuned with DPO. Every component (tokenizer, architecture, data pipeline, training loop, SFT, DPO) was written from scratch with Claude (Anthropic's AI assistant).

Model Details

Property Value
Parameters 510,985,216
Hidden size 1024
Layers 32
Query / KV heads 16 / 8 (GQA)
FFN size 3840 (SwiGLU)
Context length 8192 tokens
Vocab size 32,000
Position encoding RoPE (θ=500000)
Normalization RMSNorm (pre-norm)
Biases None
Embeddings Tied

Training

Pretraining — 10B tokens, seq_len 8192, 512K tokens/step, cosine LR 3e-4 → 3e-5 with 200-step warmup, AdamW (β₁=0.9, β₂=0.95), weight decay 0.1, 2× RTX 5060 Ti 16GB (DDP). Data mix: FineWeb-Edu, Cosmopedia, StackExchange, Wikipedia, OpenWebText, WildChat, LMSYS-Chat, UltraChat, OASST2, UltraFeedback, CodeSearchNet (Python).

SFT — 100K examples from OpenHermes-2.5, ChatML with loss masked on non-assistant tokens, 3 epochs, LR 2e-5 → 2e-6 (SFT step 1023).

DPO — 1 epoch over 62,785 preference pairs: 51,989 from UltraFeedback (tied scores dropped) and 10,796 from PKU-SafeRLHF (pairs where exactly one response is labelled safe; that one is preferred). β=0.1, LR 1e-6 → 1e-7 with 10% warmup, 64 pairs/step, frozen SFT model as reference (DPO step 981). Held-out pair accuracy rose from 0.49 to 0.71 (0.78 on the safety pairs) with no change in perplexity.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained("redptam/tyrian-500m", trust_remote_code=True,
                                             dtype=torch.bfloat16).cuda()
tokenizer = AutoTokenizer.from_pretrained("redptam/tyrian-500m", trust_remote_code=True)

prompt = tokenizer.apply_chat_template([{"role": "user", "content": "What is the capital of France?"}],
                                       tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
im_end = tokenizer.convert_tokens_to_ids("<|im_end|>")
output = model.generate(inputs["input_ids"], max_new_tokens=100, temperature=0.8, top_k=50,
                        stop_token_ids=(im_end,))
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Assistant turns end with <|im_end|> (id 5); use it as the stop token.

Serving with tyrian-serve

tyrian-serve is an OpenAI-compatible server for the Tyrian models (/v1/chat/completions, with streaming), so it can be used from Open WebUI or any OpenAI client. It runs on CUDA and falls back to CPU; on an RTX 5060 Ti this model generates about 60 tokens/s.

git clone https://github.com/redptam/tyrian-serve && cd tyrian-serve
python3 -m venv .venv
./.venv/bin/pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
./.venv/bin/hf download redptam/tyrian-500m --local-dir .modelcache/tyrian-500m
./.venv/bin/python -m uvicorn server:app --host 0.0.0.0 --port 8000
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"tyrian-500m","messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":800,"repetition_penalty":1.15}'

Use a repetition_penalty of about 1.15 and keep max_tokens at around 800 or less: answers that run past roughly 600 tokens usually start looping rather than ending. The context window is 8192 tokens.

Special Tokens

Token ID
<pad> 0
<bos> 1
<eos> 2
<unk> 3
<|im_start|> 4
<|im_end|> 5

Limitations

This is a small research model, built to learn how language models work from the ground up. It is not suitable for production use.

  • Often wrong. It writes fluent text that is frequently factually incorrect, and it states errors confidently. Do not rely on it for medical, legal, financial or other advice.
  • Not safety-tuned in practice. DPO included safety preference pairs, but red-team tests show it still does not decline harmful requests; at most it adds a warning before answering. Pretraining data includes web text and real chatbot conversations, so it can produce offensive, biased or otherwise inappropriate content.
  • Repetition. Output can loop, especially with greedy decoding; sampling with a temperature helps.
  • English only.
  • Limited long-context recall. It was trained on 8192-token sequences, but in passkey-retrieval tests it fails to recall a specific detail from more than a few hundred tokens earlier.

License

MIT

Downloads last month
542
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train redptam/tyrian-500m