File size: 11,214 Bytes
e6a9db9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a67f164
e6a9db9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9e8dbed
 
 
e6a9db9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8a65a3b
 
 
e4f9975
e6a9db9
70b1c5d
 
 
e4f9975
 
 
70b1c5d
 
 
 
 
e4f9975
 
 
 
 
 
 
 
 
 
 
 
 
 
70b1c5d
 
e4f9975
70b1c5d
e6a9db9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
---
license: mit
language:
- multilingual
tags:
- tokenizer
- bpe
- byte-level-bpe
- chatml
- routing
- moe
- robotics
- qwen
- qwen3.8
- jirack
---

# 💎 JiRack Precision is Robotics & Routing & Tool calls Tokenizer
- JiRack Precision Tokenizer . Fully compatible with Qwen3.8  models.
- Since Qwen 3.8 can understand images and video is good model for Robotics with JiRack robo tags

## NEW FOR JiRack Tokenizers family 

1. Added DeekSeek resoning tags 

2. So Do not use tool in system prompt that kills context size see what people do below ! 
- Native Tool Integration vs. Dynamic Context Injection
- Tokenizer-level tool calling yields higher precision and reliability than dynamically injecting schemas into RAG context windows, significantly optimizing both latency and token efficiency.

## SIZE
- Added Tool calls tags (
    "<|tool_call_start|>",
    "<|tool_call_end|>",
    "<|tool_result_start|>",
    "<|tool_result_end|>",
  )
  
- Size {
  Vocab size: 248191
  pad_token_id: 248044
  eos_token_id: 248046
}
- Compatible with Qwen 2.5 and DeekSeek R1 • Optimized for Code . Use resize function without adaptation, see examples below
- It needs 100k example to fully adapt routing features for RAG Routing model . So check Qwen modification tokenizer rules
- A DeekSeek R1-based tokenizer enhanced with FIM markers from Microsoft datasets and other tags for ML (
    <|fim_prefix|>,
    <|fim_middle|>,
    <f|im_suffix|>
  )
- Added Robotics & Embodiment tags (
   "<|action_start|>",
    "<|action_end|>",
    "<|trajectory_start|>",
    "<|trajectory_end|>",
    "<|joint_start|>",
    "<|joint_end|>",
    "<|sensor_start|>",
    "<|sensor_end|>",
    "<|command_start|>",
    "<|command_end|>",
    "<|state_start|>",
    "<|state_end|>",
    "<|pose|>",
    "<|velocity|>",
    "<|force|>",
    "<|torque|>",
    "<|gripper|>",
    "<|navigation|>",
    "<|obstacle|>",
    "<|task_start|>",
    "<|task_end|>",
    "<|plan_start|>",
    "<|plan_end|>",
    "<|behavior_start|>",
    "<|behavior_end|>",
    "<|skill_start|>",
    "<|skill_end|>",
    "<|motor|>",
    "<|servo|>",
    "<|imu|>",
    "<|lidar|>",
    "<|camera|>",
    "<|depth|>",
    "<|waypoint|>",
    "<|path|>",
    "<|collision|>",
    "<|grasp|>",
    "<|release|>",
    "<|homing|>",
    "<|emergency_stop|>",
    "<|calibration|>",
    "<|manipulation|>",
    "<|locomotion|>",
    "<|feedback|>",
    "<|control_loop|>",)

- Added Multi models support (
    "<|image|>",
    "<|video|>",
    "<|sound|>",
    "<|voice|>",
    "<|listening|>",
    "<|vision|>",)

- Added Human mood tags (
    "<|mood_happy|>",
    "<|mood_sad|>",
    "<|mood_angry|>",
    "<|mood_neutral|>",
  )

- Added Tool calls tags (
    "<|tool_call_start|>",
    "<|tool_call_end|>",
    "<|tool_result_start|>",
    "<|tool_result_end|>",
  )


- Added RAG routing tags for RAG MoE Systems (
    "__SCIENCE__",
    "__CODING__",
    "__STOCK_EXCHANGE__",
    "__MEDICINE__",
    "__GOVERNMENT__",
    "__NEWS__",
    "__GENERAL__",
    "__MATERIAL_SCIENCE__",
    "__ELECTRONICS__",
    "__MICROELECTRONICS__",
    "__ENGINEERING__",
    "__ROBOTICS__",
    "__ENERGY__",
    "__AUTOMOTIVE__",
    "__AVIATION__",
    "__MATH__",
    "__PYTHON__",
    "__C__",
    "__CPP__",
    "__C_SHARP__",
    "__JAVA__",
    "__JAVASCRIPT__",
    "__TYPESCRIPT__",
    "__RUST__",
    "__GO__",
    "__RUBY__",
    "__PHP__",
    "__SWIFT__",
    "__KOTLIN__",
    "__BASH__",
    "__SQL__",
    "__ASSEMBLY__",
    "__PHILOSOPHY__",
    "__LITERATURE__",
    "__SOCIOLOGY__",
    "__PSYCHOLOGY__",
    "__POLITICAL_SCIENCE__",
    "__CULTURAL_STUDIES__",
    "__ETHNOGRAPHY__",
    "__HUMAN_RIGHTS__",
    "__COMPLIANCE__",
    "__MILITARY__",
    "__BANKING__",
    "__OIL_INDUSTRY__",
    "__LIGHT_INDUSTRY__",
    "__NATURE__",
    "__OCEAN__",
    "__SPORT__",
    "__CULINARY__",
    "__TRAVEL__",
    "__HOBBY__"
  )
    
- Fully compatible with Microsoft BigCode datasets including The Stack, StarCoder, and NextCoder.
- Enables efficient training on large-scale coding data for superior code generation and understanding.


**Inventor:** Konstantin Vladimirovich Grabko  
**Organization:** CMS Manhattan JiRack Technology  
**Official Site:** [www.cmsmanhattan.com](http://www.cmsmanhattan.com)
Designed for Banking and Fintech Institutions

**Banks and Fintech**  JiRack Architecture: Build Sovereign Financial Models from Scratch
- Leveraging the JiRack Tokenizer and our Open Dataset, we enable financial institutions to develop secure, internal AI models from the ground up. This approach ensures maximum data privacy and model sovereignty for high-stakes banking operations.
- There is fix price for FinTech
- I recommend initializing the model with a 4K context window for initial stability, followed by scaling to 8K context using specialized JiRack 8K datasets. This two-stage approach ensures robust positional encoding before extending the model's long-range dependency.


**JiRack Corp Tokenizer  solution**
- Use JiRack models with trusted, high-quality coding datasets while maintaining full control over your code and data privacy.
- Excellent fit for Banks, Fintech companies, and any organization that requires strict data confidentiality and security.
- Update JiRack model for corp privacy coding. 

**JiRack Tokenizer Subcription**
- All subscribed members will receive regular tokenizer updates optimized for the latest high-quality coding datasets.

### Open Robot platform 

- **Tiangong** : https://english.www.gov.cn/english.www.gov.cn/news/202411/13/content_WS673406e2c6d0868f4e8ece33.html
- **Unitree g1**  https://a.co/d/0e4A8YVc
- **LimX Oli**  https://www.limxdynamics.com/en/products/oli?channel=option_google_advertising__c-
- **ubtrobot** https://www.ubtrobot.com/en/
- **x-humanoid** https://www.x-humanoid.com/detail/hskw.html
- **Drone**  https://arrowair.com/quiver/

## ROBOS HUMANOID Project

- https://github.com/cmsmanhattan/ROBOS
  

### Robot dataset example

```bash
{"messages":[{"role":"system","content":"You are the robot's brain. Use tags and . ROBOTICS"},{"role":"user","content":"Scene: a table in front of me, with a red mug on it. Command: pick up the mug."},{"role":"assistant","content":"reach(x=0.4,y=0.0,z=0.3)"}]}
{"messages":[{"role":"system","content":"You are the robot's brain. Use tags and .ROBOTICS"},{"role":"user","content":"Scene: hand near the mug, distance 2 cm. Command: grasp."},{"role":"assistant","content":"grasp(force=60)"}]}
{"messages":[{"role":"system","content":"You are the robot's brain. Use tags and . ROBOTICS"},{"role":"user","content":"Scene: mug in hand, stable. Command: lift 20 cm."},{"role":"assistant","content":"lift(z=0.2)"}]}

```

```bash
{
 "messages": [
 {
 "role": "system",
 "content": "You are a robot. Use tags and . ROBOTICS"
 },
 {
 "role": "user",
 "content": "Scene: a table in front of me, with a red mug on it. Command: pick up the mug."
 },
 {
 "role": "assistant",
 "content": "reach(x=0.4,y=0.0,z=0.3)"
 }
]
}
```
```

### Key Features

- **Algorithm**: Byte-Level BPE
- **Vocabulary Size**: **128,000** tokens — excellent balance between precision and efficiency
- **Multilingual & Technical Strength**: Optimized for English, Russian, code, scientific literature, and technical documentation
- **Domain Specialization**: Strong performance on programming languages, engineering, robotics, and scientific texts


### Tool calling 
- Here's a concrete example of the Qwen 2.5 / Hermes-style tool-calling format: https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1
  
- System message (defines available tools):
  
```bash
<|im_start|>system
You are a helpful assistant with access to the following functions. Use them if required:
<tools>
{
"type": "function",
"function": {"name": "get_weather", "description": "Get current weather for a location",
"parameters":
  {"type": "object", "properties":
   {
    "location":{"type": "string", "description": "City name"},
    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]
   }
  }
}
</tools>
For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>
<|im_end|>
```

- User message:

```bash
<|im_start|>user What's the weather like in Tokyo right now?<|im_end|>
```

- Assistant message (model emits the tool call instead of a normal answer):
  
```bash  
<|im_start|>assistant
 <tool_call>
   {"name": "get_weather", "arguments": {"location": "Tokyo", "unit": "celsius"}}
 </tool_call>
<|im_end|>
```

- Tool result message (your agent code executes the function and feeds the result back):
  
```bash
<|im_start|>tool
 <tool_response>
  {"temperature": 22, "condition": "clear sky", "humidity": 48}
 </tool_response>
<|im_end|>
```

- Final assistant message (model now answers using the result):
  
```bash
<|im_start|>assistant It's currently 22°C and clear in Tokyo, with 48% humidity.<|im_end|>
```

```bash
That's the exact skeleton your Hermes dataset entries need to be mapped into 
— system prompt with <tools>...</tools>, <tool_call> in assistant turns, <tool_response> in tool turns. 
If you apply Qwen 2.5's tokenizer chat template via apply_chat_template() with a tools=[...] argument, 
it generates this structure automatically — you don't have to hand-write the tags yourself for training data prep, 
as long as your data is in the standard role/content/tool_calls message-list format first
```

### Special Tokens Support

- Full **Qwen2 compatible format** dialogue format 
- FIM (Fill-in-the-Middle) support for code generation
- Rich set of domain routing tokens (`__CODING__`, `__PYTHON__`, `__ROBOTICS__`, `__SCIENCE__`, etc.)
- Extended robotics and control tokens


### CMS Manhattan Service & Support 

- Jirack patent guards your technology for competitors 
- Redesign Llama , Qwen , Gemma  to Ternary model
- Re-tain and replace embeddings for  Llama , Qwen  , Gemma  to extend langeages to 347
- Accelerate inference via high compression 256K tokenizer and replace multiplication with sum operations via Ternary weights
  


**Install for Llamma compatible models in your chat script**

- from transformers import AutoModelForCausalLM

- model = AutoModelForCausalLM.from_pretrained(your model)

- # The Must !
- model.resize_token_embeddings(len(tokenizer.tokenizer))   # или просто len(tokenizer.tokenizer)

- print("New Embedding size for you chat script:", model.get_input_embeddings().weight.shape[0])

- # Tesr Tokenizer size !
(venv_ji) root@jirack2:# python -c '
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("./QwenRoboticsTokenizer")
print("Vocab size:", len(tok))
print("pad_token_id:", tok.pad_token_id)
print("eos_token_id:", tok.eos_token_id)
'

Vocab size: 151778
pad_token_id: 151643
eos_token_id: 151645

## 📧 Contact & Licensing

For joint ventures, hardware integration, or licensing inquiries:

- **Email:** grabko@cmsmanhattan.com
- **Phone:** +1 (516) 777-0945
- **Location:** New York, USA

## 📧 Copyright 2026  CMS Manhattan . All rights reserved

## License
MIT License