Monster-Code commited on
Commit
fbf7a3d
·
verified ·
1 Parent(s): f1d23a1

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +175 -2
README.md CHANGED
@@ -1,5 +1,178 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
4
- # Boomslang: a 3B parameter reasoning model based off Qwen-2.5-3B.
5
- ![1789911910f175](https://cdn-uploads.huggingface.co/production/uploads/6a5f8c87a38cac087c2c6c05/Ag4eWdDI71YPtuwsP7Z5e.png)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-3B-Instruct
4
+ language:
5
+ - en
6
+ tags:
7
+ - reasoning
8
+ - math
9
+ - chain-of-thought
10
+ - qwen2.5
11
+ - gguf
12
+ - text-generation
13
+ - conversational
14
+ pipeline_tag: text-generation
15
+ library_name: transformers
16
  ---
17
+
18
+ <div align="center">
19
+
20
+ # 🐍 Boomslang (3B Reasoning & Math Engine)
21
+
22
+ **A compact, high-efficiency 3-billion-parameter model fine-tuned for deep chain-of-thought mathematical reasoning, logic, and algebra—without sacrificing natural conversational ability.**
23
+
24
+ [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Creator-Monster--Code-blue)](https://huggingface.co/Monster-Code)
25
+ [![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](https://opensource.org/licenses/Apache-2.0)
26
+ [![GGUF Available](https://img.shields.io/badge/GGUF-Included-purple.svg)](https://huggingface.co/Monster-Code/Boomslang/blob/main/boomslang-3b-qwen.gguf)
27
+
28
+ ---
29
+
30
+ ### ❤️ If you find Boomslang useful, please hit the **Like** button at the top of this page and [Follow @Monster-Code](https://huggingface.co/Monster-Code) for more open-weights AI releases!
31
+
32
+ ---
33
+
34
+ </div>
35
+
36
+ ## 💡 What is Boomslang?
37
+
38
+ Most small models (1B–3B parameters) struggle with two extremes: they are either polite chatbots that completely hallucinate basic arithmetic, or narrow math models that forget how to hold a conversation and start writing unprompted proofs when you simply say "Hi."
39
+
40
+ **Boomslang** was trained to bridge that gap.
41
+
42
+ Starting from the strong foundation of **`Qwen/Qwen2.5-3B-Instruct`**, Boomslang was post-trained on an **NVIDIA RTX PRO 6000 Blackwell** across a curated ~17,500-sample reasoning mixture:
43
+ 1. **DeepSeek-R1 Distilled Proofs (`open-r1/OpenR1-Math-220k`):** Teaches the network an internal self-reflection loop (`<think> ... </think>`) to break down complex algebraic expressions, geometry, and multi-step deduction before committing to an answer.
44
+ 2. **Step-by-Step Arithmetic Rigor (`openai/gsm8k`):** Calibrates attention heads on strict order-of-operations arithmetic and unambiguous answer derivation.
45
+
46
+ The result is an edge-friendly 3B model that works through tricky algebra and word puzzles methodically, but still greets you warmly and follows instructions when you just want to talk.
47
+
48
+ ---
49
+
50
+ ## 📦 What's Inside This Repository?
51
+
52
+ * **Single Standalone `model.safetensors`:** No multi-part file splits. The entire 3-billion parameter model is packed into a single, clean ~6.1 GB file.
53
+ * **Pre-Converted GGUF (`boomslang-3b-qwen.gguf`):** Directly ready for **Ollama**, **LM Studio**, and **llama.cpp** on your local machine (MacBook, laptop, or home GPU).
54
+ * Full configuration and tokenizer files for immediate `transformers` plug-and-play.
55
+
56
+ ---
57
+
58
+ ## ⚡ Quickstart: Python & Transformers
59
+
60
+ ```python
61
+ import torch
62
+ from transformers import AutoTokenizer, AutoModelForCausalLM, TextStreamer
63
+
64
+ MODEL_ID = "Monster-Code/Boomslang"
65
+
66
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
67
+ model = AutoModelForCausalLM.from_pretrained(
68
+ MODEL_ID,
69
+ torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
70
+ device_map="auto"
71
+ )
72
+
73
+ # Standard ChatML format
74
+ messages = [
75
+ {
76
+ "role": "system",
77
+ "content": "You are Boomslang, a helpful and precise reasoning assistant. Reason step-by-step before providing your final answer."
78
+ },
79
+ {
80
+ "role": "user",
81
+ "content": "A box has 36 chocolates. Tom eats 1/4 of them, and Sarah eats 9. How many chocolates are left in the box?"
82
+ }
83
+ ]
84
+
85
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
86
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
87
+ streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
88
+
89
+ with torch.no_grad():
90
+ model.generate(
91
+ **inputs,
92
+ streamer=streamer,
93
+ max_new_tokens=1024,
94
+ temperature=0.6,
95
+ top_p=0.9
96
+ )
97
+ ```
98
+
99
+ ---
100
+
101
+ ## 💻 Running Locally with Ollama / LM Studio
102
+
103
+ ### 1. Using Ollama:
104
+ Create a `Modelfile` in the same directory as the downloaded `.gguf`:
105
+ ```dockerfile
106
+ FROM ./boomslang-3b-qwen.gguf
107
+
108
+ TEMPLATE """{{ if .System }}<|im_start|>system
109
+ {{ .System }}<|im_end|>
110
+ {{ end }}{{ if .Prompt }}<|im_start|>user
111
+ {{ .Prompt }}<|im_end|>
112
+ {{ end }}<|im_start|>assistant
113
+ {{ .Response }}<|im_end|>"""
114
+
115
+ PARAMETER stop "<|im_start|>"
116
+ PARAMETER stop "<|im_end|>"
117
+ PARAMETER temperature 0.6
118
+ ```
119
+
120
+ Build and run:
121
+ ```bash
122
+ ollama create boomslang -f Modelfile
123
+ ollama run boomslang
124
+ ```
125
+
126
+ ### 2. Using LM Studio:
127
+ 1. Download **`boomslang-3b-qwen.gguf`** from the [Files and versions tab](https://huggingface.co/Monster-Code/Boomslang/tree/main).
128
+ 2. Drop it into your local LM Studio models folder.
129
+ 3. Start chatting with GPU offloading enabled!
130
+
131
+ ---
132
+
133
+ ## 🧠 Sample Interaction
134
+
135
+ **User:**
136
+ > *"Solve for x: 5x - 10 = 25. Show your work."*
137
+
138
+ **Boomslang:**
139
+ ```text
140
+ <think>
141
+ The equation is 5x - 10 = 25.
142
+ Our goal is to isolate the variable x on one side.
143
+
144
+ Step 1: Add 10 to both sides to cancel out the -10 constant:
145
+ 5x - 10 + 10 = 25 + 10
146
+ 5x = 35
147
+
148
+ Step 2: Divide both sides by the coefficient 5:
149
+ 5x / 5 = 35 / 5
150
+ x = 7
151
+
152
+ Verification:
153
+ Substitute x = 7 into the original expression:
154
+ 5(7) - 10 = 35 - 10 = 25.
155
+ Both sides match.
156
+ </think>
157
+
158
+ The answer is x = 7.
159
+ ```
160
+
161
+ ---
162
+
163
+ ## 🛠️ Training Details
164
+
165
+ * **Base Architecture:** Qwen2.5 3B (Decoder-only Transformer)
166
+ * **Hardware:** NVIDIA RTX PRO 6000 Blackwell Server Edition
167
+ * **Precision:** BF16 Native Mixed Precision with Fused AdamW
168
+ * **Effective Batch Size:** 32 (8 per device × 4 gradient accumulation steps)
169
+ * **Learning Rate:** 1.5e-4 with dynamic linear warmup
170
+ * **Label Masking:** Dynamic batch padding via `DataCollatorForSeq2Seq` with `-100` masking to guarantee loss is never calculated on padding noise
171
+
172
+ ---
173
+
174
+ ## 🤝 Community & Support
175
+
176
+ * 👤 **Creator:** [Monster-Code](https://huggingface.co/Monster-Code)
177
+ * 💬 Have suggestions, evaluation runs, or dataset ideas? Leave a note in the **Discussions** tab!
178
+ * ⭐ **If Boomslang helps your workflow, please consider starring/liking the repository!**