Instructions to use redptam/tyrian-500m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use redptam/tyrian-500m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="redptam/tyrian-500m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("redptam/tyrian-500m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use redptam/tyrian-500m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "redptam/tyrian-500m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-500m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/redptam/tyrian-500m
- SGLang
How to use redptam/tyrian-500m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "redptam/tyrian-500m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-500m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "redptam/tyrian-500m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-500m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use redptam/tyrian-500m with Docker Model Runner:
docker model run hf.co/redptam/tyrian-500m
Tyrian 500M
A 511M parameter decoder-only language model built from scratch in PyTorch, then supervised fine-tuned for chat (ChatML) and preference-tuned with DPO. Every component (tokenizer, architecture, data pipeline, training loop, SFT, DPO) was written from scratch with Claude (Anthropic's AI assistant).
Model Details
| Property | Value |
|---|---|
| Parameters | 510,985,216 |
| Hidden size | 1024 |
| Layers | 32 |
| Query / KV heads | 16 / 8 (GQA) |
| FFN size | 3840 (SwiGLU) |
| Context length | 8192 tokens |
| Vocab size | 32,000 |
| Position encoding | RoPE (θ=500000) |
| Normalization | RMSNorm (pre-norm) |
| Biases | None |
| Embeddings | Tied |
Training
Pretraining — 10B tokens, seq_len 8192, 512K tokens/step, cosine LR 3e-4 → 3e-5 with 200-step warmup, AdamW (β₁=0.9, β₂=0.95), weight decay 0.1, 2× RTX 5060 Ti 16GB (DDP). Data mix: FineWeb-Edu, Cosmopedia, StackExchange, Wikipedia, OpenWebText, WildChat, LMSYS-Chat, UltraChat, OASST2, UltraFeedback, CodeSearchNet (Python).
SFT — 100K examples from OpenHermes-2.5, ChatML with loss masked on non-assistant tokens, 3 epochs, LR 2e-5 → 2e-6 (SFT step 1023).
DPO — 1 epoch over 62,785 preference pairs: 51,989 from UltraFeedback (tied scores dropped) and 10,796 from PKU-SafeRLHF (pairs where exactly one response is labelled safe; that one is preferred). β=0.1, LR 1e-6 → 1e-7 with 10% warmup, 64 pairs/step, frozen SFT model as reference (DPO step 981). Held-out pair accuracy rose from 0.49 to 0.71 (0.78 on the safety pairs) with no change in perplexity.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained("redptam/tyrian-500m", trust_remote_code=True,
dtype=torch.bfloat16).cuda()
tokenizer = AutoTokenizer.from_pretrained("redptam/tyrian-500m", trust_remote_code=True)
prompt = tokenizer.apply_chat_template([{"role": "user", "content": "What is the capital of France?"}],
tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
im_end = tokenizer.convert_tokens_to_ids("<|im_end|>")
output = model.generate(inputs["input_ids"], max_new_tokens=100, temperature=0.8, top_k=50,
stop_token_ids=(im_end,))
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Assistant turns end with <|im_end|> (id 5); use it as the stop token.
Serving with tyrian-serve
tyrian-serve is an OpenAI-compatible server for the Tyrian
models (/v1/chat/completions, with streaming), so it can be used from Open WebUI or any OpenAI client.
It runs on CUDA and falls back to CPU; on an RTX 5060 Ti this model generates about 60 tokens/s.
git clone https://github.com/redptam/tyrian-serve && cd tyrian-serve
python3 -m venv .venv
./.venv/bin/pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
./.venv/bin/hf download redptam/tyrian-500m --local-dir .modelcache/tyrian-500m
./.venv/bin/python -m uvicorn server:app --host 0.0.0.0 --port 8000
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"tyrian-500m","messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":800,"repetition_penalty":1.15}'
Use a repetition_penalty of about 1.15 and keep max_tokens at around 800 or less: answers that run past
roughly 600 tokens usually start looping rather than ending. The context window is 8192 tokens.
Special Tokens
| Token | ID |
|---|---|
<pad> |
0 |
<bos> |
1 |
<eos> |
2 |
<unk> |
3 |
<|im_start|> |
4 |
<|im_end|> |
5 |
Limitations
This is a small research model, built to learn how language models work from the ground up. It is not suitable for production use.
- Often wrong. It writes fluent text that is frequently factually incorrect, and it states errors confidently. Do not rely on it for medical, legal, financial or other advice.
- Not safety-tuned in practice. DPO included safety preference pairs, but red-team tests show it still does not decline harmful requests; at most it adds a warning before answering. Pretraining data includes web text and real chatbot conversations, so it can produce offensive, biased or otherwise inappropriate content.
- Repetition. Output can loop, especially with greedy decoding; sampling with a temperature helps.
- English only.
- Limited long-context recall. It was trained on 8192-token sequences, but in passkey-retrieval tests it fails to recall a specific detail from more than a few hundred tokens earlier.
License
MIT
- Downloads last month
- 542