File size: 1,633 Bytes
759d613
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
---

language:
- en
license: mit
library_name: transformers
pipeline_tag: text-generation
tags:
- tiny-models
- custom-architecture
- story-generation
- experimental
---


# Spin-80k

**Spin-80k** is a lightweight, 80k-parameter decoder-only language model built from scratch by **Quantech** to demonstrate custom Transformer architecture 

---

## Model Specifications

* **Organization:** Quantech
* **Architecture:** Custom Decoder-only Transformer
* **Total Parameters:** ~80,112
* **Layers:** 2
* **Hidden Dimension ($d_{\text{model}}$):** 48

* **Attention Heads:** 4

* **Feed-Forward Dimension ($d_{\text{ff}}$):** 128
* **Positional Encoding:** Rotary Position Embeddings (RoPE)
* **Normalization:** RMSNorm ($\epsilon = 10^{-5}$)
* **Activation:** SwiGLU
* **Vocabulary:** 512 Byte-Pair Encoding (BPE) tokens
* **Context Length:** 256 tokens

---

## Quickstart

```python

import torch

from transformers import AutoModelForCausalLM, AutoTokenizer



repo_id = "Quantech/spin-80k"



# Load Tokenizer & Model

tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)

model = AutoModelForCausalLM.from_pretrained(repo_id, trust_remote_code=True)

model.eval()



# ChatML Format

prompt = "<|im_start|>user\nWrite a short story about a dog.<|im_end|>\n<|im_start|>assistant\n"

inputs = tokenizer(prompt, return_tensors="pt")



with torch.no_grad():

    outputs = model.generate(

        **inputs,

        max_new_tokens=50,

        temperature=0.7,

        do_sample=True,

        pad_token_id=tokenizer.eos_token_id

    )



print(tokenizer.decode(outputs[0]))