--- license: apache-2.0 datasets: - HuggingFaceFW/fineweb-edu - HuggingFaceTB/finemath - HuggingFaceTB/smollm-corpus --- ![Banner](https://cdn-uploads.huggingface.co/production/uploads/685ea8ff7b4139b6845ce395/0QdBKT5iKdrEy3iwLosAH.png) # DynamicMind-Mini DynamicMind-Mini is a decoder-only model trained on [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), [SmolLM-Corpus](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) and [FineMath](https://huggingface.co/datasets/HuggingFaceTB/finemath) The model has about 8.9M parameters and its based on [MiniBananaMind-v4-9M](https://huggingface.co/BananaMind/MiniBananaMind-v4-9M) architecture with a custom 8k-token byte-level BPE tokenizer with digit-aware tokenization. ## Model Details | Field | Value | |---|---:| | Parameters | 8,884,992 | | Architecture | Custom Llama-style decoder | | Layers | 9 | | Hidden size | 256 | | Intermediate size | 768 | | Attention heads | 8 | | KV heads | 2 | | Vocabulary size | 8,192 | | Context length | 1,024 | | Embeddings | Tied input/output embeddings | | Weight format | safetensors | | Base checkpoint | MiniBananaMind-v3-9M | | Continued pretraining checkpoint | Cosmopedia-v2 step 2,613 | ## Tokenizer DynamicMind-Mini uses the same digit-aware 8k tokenizer as [MiniBananaMind-v4-9M](https://huggingface.co/BananaMind/MiniBananaMind-v4-9M). Digits are kept as separate tokens so numbers do not collapse into large number tokens during tokenization. Digit IDs: | Token | ID | |---|---:| | `1` | 9 | | `2` | 10 | | `3` | 11 | | `4` | 12 | | `5` | 13 | | `6` | 14 | | `7` | 15 | | `8` | 16 | | `9` | 17 | | `0` | 18 | ## Usage This model uses custom architecture code, so load it with `trust_remote_code=True`. Install dependencies: ```bash pip install -U transformers safetensors torch ``` Run inference: ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM model_id = "DedeProGames/DynamicMind-Mini" tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, torch_dtype=torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float16, ).cuda().eval() prompt = "The meaning of life is " input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device) with torch.no_grad(): output = model.generate( input_ids=input_ids, max_new_tokens=64, do_sample=False, repetition_penalty=1.1, pad_token_id=tokenizer.eos_token_id, eos_token_id=tokenizer.eos_token_id, ) print(tokenizer.decode(output[0], skip_special_tokens=True)) ```