devoppro commited on
Commit
ebbe37b
·
verified ·
1 Parent(s): 6608932

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +52 -0
README.md ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - custom-architecture
9
+ - rope
10
+ - gqa
11
+ - swiglu
12
+ - rmsnorm
13
+ - safetensors
14
+ ---
15
+
16
+ # FastLLM (150M) — Modern Causal Language Model
17
+
18
+ **FastLLM** is a ~150M parameter, decoder-only causal language model built completely from scratch in PyTorch and fully integrated with Hugging Face `transformers`. It incorporates state-of-the-art LLM architectural choices—**Grouped-Query Attention (GQA)**, **SwiGLU MLPs**, **RMSNorm**, and **Rotary Position Embeddings (RoPE)**—and natively saves weights in the zero-copy **Safetensors** format.
19
+
20
+ ---
21
+
22
+ ## Model Details
23
+
24
+ * **Developed by:** devoppro
25
+ * **Model Type:** Decoder-only Causal Language Model
26
+ * **Architecture:** Custom Transformer (`ModernLLMForCausalLM`)
27
+ * **Parameter Count:** ~150,000,000 (150M)
28
+ * **Tokenizer:** Qwen 2.5 BPE Vocabulary (`vocab_size`: 151,936)
29
+ * **Precision:** Mixed Precision (`FP16`)
30
+ * **Storage Format:** `.safetensors`
31
+ * **Repository:** `devoppro/FastLLM`
32
+
33
+ ---
34
+
35
+ ## Architectural Specifications
36
+
37
+ | Parameter | Configuration |
38
+ | :--- | :--- |
39
+ | **Hidden Size ($d_{\text{model}}$)** | 768 |
40
+ | **Intermediate Size (SwiGLU)** | 2048 |
41
+ | **Hidden Layers** | 12 |
42
+ | **Query Heads** | 12 |
43
+ | **Key/Value Heads (GQA)** | 4 (3:1 Query-to-KV ratio) |
44
+ | **Max Context Length** | 2048 tokens |
45
+ | **Normalization** | RMSNorm ($\epsilon = 10^{-6}$) |
46
+ | **Positional Embedding** | Rotary Embeddings (RoPE, $\theta = 1000000.0$) |
47
+
48
+ ---
49
+
50
+ ## Training Data Mixture
51
+
52
+ The model was pre-trained using dynamic stream interleaving across four high-quality datasets: