Instructions to use redptam/tyrian-75m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use redptam/tyrian-75m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="redptam/tyrian-75m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("redptam/tyrian-75m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use redptam/tyrian-75m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "redptam/tyrian-75m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-75m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/redptam/tyrian-75m
- SGLang
How to use redptam/tyrian-75m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "redptam/tyrian-75m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-75m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "redptam/tyrian-75m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "redptam/tyrian-75m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use redptam/tyrian-75m with Docker Model Runner:
docker model run hf.co/redptam/tyrian-75m
Tyrian 75M
A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant).
Model Details
| Property | Value |
|---|---|
| Parameters | 74,920,704 (~75M) |
| Architecture | Decoder-only transformer |
| Hidden size | 768 |
| Layers | 8 |
| Query heads | 12 (GQA) |
| KV heads | 4 (GQA) |
| FFN size | 2048 (SwiGLU) |
| Context length | 2048 tokens |
| Vocab size | 32,000 |
| Normalization | RMSNorm (pre-norm) |
| Position encoding | RoPE (θ=10000) |
| Attention | Flash Attention (SDPA) |
| FFN activation | SwiGLU |
| Biases | None |
| Embeddings | Tied (input = output) |
Training
Pretraining
- ~17.8B tokens of English web text
- Data mix: FineWeb-Edu, Cosmopedia, StackExchange, Wikipedia, OpenWebText, WildChat, UltraChat, OASST2
- Custom BPE tokenizer (32K vocab, ChatML format)
- Cosine LR schedule: 3e-4 → 3e-5 with 2000-step warmup
- AdamW (β₁=0.9, β₂=0.95), weight decay 0.1
- Batch: 512K tokens/step
- Hardware: 2× RTX 5060 Ti 16GB, DDP
SFT Fine-tuning
- 100K examples from OpenHermes-2.5
- ChatML format with loss masking on user/system tokens
- LR: 2e-5 → 2e-6, 3 epochs
Benchmarks
Evaluated with log-likelihood scoring (no few-shot):
| Task | Score | Random |
|---|---|---|
| HellaSwag | 27.8% | 25% |
| PIQA | 61.0% | 50% |
| ARC-Easy | 39.3% | 25% |
| ARC-Challenge | 25.4% | 25% |
| WinoGrande | 51.7% | 50% |
Comparable to GPT-2 (117M) at 0.64× the parameter count.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"redptam/tyrian-75m",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).cuda()
tokenizer = AutoTokenizer.from_pretrained("redptam/tyrian-75m", trust_remote_code=True)
# Chat (ChatML format)
prompt = "<|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=100, temperature=0.8, top_k=50,
stop_token_ids=(tokenizer.convert_tokens_to_ids("<|im_end|>"),))
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))
Serving with tyrian-serve
tyrian-serve is an OpenAI-compatible server for the Tyrian
models (/v1/chat/completions, with streaming), so it can be used from Open WebUI or any OpenAI client.
It runs on CUDA and falls back to CPU.
git clone https://github.com/redptam/tyrian-serve && cd tyrian-serve
python3 -m venv .venv
./.venv/bin/pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
./.venv/bin/hf download redptam/tyrian-75m --local-dir .modelcache/tyrian-75m
MODEL_PATH=.modelcache/tyrian-75m MODEL_NAME=tyrian-75m ./.venv/bin/python -m uvicorn server:app --host 0.0.0.0 --port 8000
curl -s http://localhost:8000/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"tyrian-75m","messages":[{"role":"user","content":"What is the capital of France?"}],"repeat_penalty":1.2}'
Set a repeat_penalty of about 1.2–1.3 in requests (or in the client's model settings); without one the 75M
loops badly. The context window is 2048 tokens.
Special Tokens
| Token | ID |
|---|---|
<pad> |
0 |
<bos> |
1 |
<eos> |
2 |
<unk> |
3 |
<|im_start|> |
4 |
<|im_end|> |
5 |
Limitations
This is a small research model, built to learn how language models work from the ground up. It is not suitable for production use.
- Often wrong. It writes fluent text that is frequently factually incorrect, and it states errors confidently. Do not rely on it for medical, legal, financial or other advice.
- Not safety-tuned. It has had supervised fine-tuning only, with no preference or safety tuning, so it will not reliably decline harmful or inappropriate requests. Pretraining data includes web text and real chatbot conversations, so it can produce offensive, biased or otherwise inappropriate content.
- Repetition. Output can loop, especially with greedy decoding; sampling with a temperature helps.
- English only.
- Weak reasoning. Benchmark scores are close to chance on ARC-Challenge and WinoGrande (see above), and the context window is 2048 tokens.
License
MIT
- Downloads last month
- 695