| --- |
| license: apache-2.0 |
| datasets: |
| - HuggingFaceFW/fineweb-edu |
| - HuggingFaceTB/finemath |
| - HuggingFaceTB/smollm-corpus |
| --- |
| |
|  |
|
|
| # DynamicMind-Mini |
|
|
| DynamicMind-Mini is a decoder-only model trained on [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) and [FineMath](https://huggingface.co/datasets/HuggingFaceTB/finemath) |
|
|
| The model has about 8.9M parameters and its based on [MiniBananaMind-v4-9M](https://huggingface.co/BananaMind/MiniBananaMind-v4-9M) architecture with a custom 8k-token byte-level BPE tokenizer with digit-aware tokenization. |
|
|
| ## Model Details |
|
|
| | Field | Value | |
| |---|---:| |
| | Parameters | 8,884,992 | |
| | Architecture | Custom Llama-style decoder | |
| | Layers | 9 | |
| | Hidden size | 256 | |
| | Intermediate size | 768 | |
| | Attention heads | 8 | |
| | KV heads | 2 | |
| | Vocabulary size | 8,192 | |
| | Context length | 1,024 | |
| | Embeddings | Tied input/output embeddings | |
| | Weight format | safetensors | |
| | Base checkpoint | MiniBananaMind-v3-9M | |
| | Continued pretraining checkpoint | Cosmopedia-v2 step 2,613 | |
|
|
| ## Tokenizer |
|
|
| DynamicMind-Mini uses the same digit-aware 8k tokenizer as [MiniBananaMind-v4-9M](https://huggingface.co/BananaMind/MiniBananaMind-v4-9M). |
|
|
| Digits are kept as separate tokens so numbers do not collapse into large number tokens during tokenization. |
|
|
| Digit IDs: |
|
|
| | Token | ID | |
| |---|---:| |
| | `1` | 9 | |
| | `2` | 10 | |
| | `3` | 11 | |
| | `4` | 12 | |
| | `5` | 13 | |
| | `6` | 14 | |
| | `7` | 15 | |
| | `8` | 16 | |
| | `9` | 17 | |
| | `0` | 18 | |
|
|
| ## Usage |
|
|
| This model uses custom architecture code, so load it with `trust_remote_code=True`. |
|
|
| Install dependencies: |
|
|
| ```bash |
| pip install -U transformers safetensors torch |
| ``` |
|
|
| Run inference: |
|
|
| ```python |
| import torch |
| from transformers import AutoTokenizer, AutoModelForCausalLM |
| |
| model_id = "DedeProGames/DynamicMind-Mini" |
| |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) |
| |
| model = AutoModelForCausalLM.from_pretrained( |
| model_id, |
| trust_remote_code=True, |
| torch_dtype=torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float16, |
| ).cuda().eval() |
| |
| prompt = "The meaning of life is " |
| input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device) |
| |
| with torch.no_grad(): |
| output = model.generate( |
| input_ids=input_ids, |
| max_new_tokens=64, |
| do_sample=False, |
| repetition_penalty=1.1, |
| pad_token_id=tokenizer.eos_token_id, |
| eos_token_id=tokenizer.eos_token_id, |
| ) |
| |
| print(tokenizer.decode(output[0], skip_special_tokens=True)) |
| ``` |