--- license: mit tags: - small-language-model - tiny-language-model - word-generation - distillation - synthetic-data datasets: - Harley-ml/es-en-words pipeline_tag: text-generation --- # DistilWord2-23k **DistilWord2-23k** is the second-generation ultra-compact language model (~23k parameters) in the DistilWord series. Trained on **12,000 clean, low-entropy synthetic words** generated by [Harley-ml/LargeWord-1.5M](https://huggingface.co/Harley-ml/LargeWord-1.5M), it achieves state-of-the-art sample efficiency and morphology for its size class. ## Benchmark Progression | Model | Parameters | Training Distribution | Val Loss | Val PPL | Signal PSNR | Behavior | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | **MicroWord-23k (Original)** | ~23k | 750k Natural Words | 6.4105 | 608.18 | 2.57 dB | Attention collapsed (`wwww...`, `zzzz...`) | | **DistilWord-23k (v1)** | ~23k | 3.8k Synthetic (`TinyWord2`) | 3.0633 | 21.40 | 4.37 dB | Stutter eliminated; basic morphology | | **DistilWord2-23k (Ours)** | **~23k** | **12k Synthetic (`LargeWord-1.5M`)** | **2.5415** | **12.70** | **4.50 dB** | **Real morphemes (`afields`, `appers`, `zrings`)** | --- ## Qualitative Generations | Prompt | Original MicroWord-23k | DistilWord2-23k (Ours) | Learned Morphology | | :--- | :--- | :--- | :--- | | `a` | `a` (Stalled) | **`afields`** | Complete dictionary compound word (*a* + *field* + *-s*) | | `app` | `appco` | **`appers`** | Agent noun pluralization (*app* + *-er* + *-s*) | | `z` | `zzzzx's` (Stutter loop) | **`zrings`** | Complex inflectional cluster (*-ing* + *-s*) | | `dis` | `dis` | **`disers`** | Morpheme suffix attachment | | `el` | `elel` (Repeat loop) | **`elys`** | Adverbial suffix (*-ly* + *-s*) | | `sub` | `sub` | **`subs`** | Standard plural inflection | --- ## Quick Inference ```python import torch from transformers import AutoModelForCausalLM, PreTrainedTokenizerFast repo_id = "Useruser2statsaltalt/DistilWord2-23k" tokenizer = PreTrainedTokenizerFast.from_pretrained(repo_id) model = AutoModelForCausalLM.from_pretrained(repo_id) model.eval() prompt = "app" bos = tokenizer.bos_token or "" inputs = tokenizer(bos + prompt, return_tensors="pt", add_special_tokens=False) with torch.inference_mode(): outputs = model.generate( **inputs, max_new_tokens=12, do_sample=True, temperature=0.80, top_p=0.90, top_k=35, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.pad_token_id, ) prompt_len = inputs["input_ids"].shape[-1] completion = tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True) print(f"Generated: {prompt + completion}") ## Architecture Specifications Architecture: Qwen Causal LM (Micro-scale) Hidden Layers: 1 Hidden Size: 16 Attention Heads: 1 Intermediate (SwiGLU): 56 Tied Embeddings: True Unique Parameters: ~23,024