DistilWord-23k

DistilWord-23k is an ultra-compact language model trained on 3.8k synthetic words from Harley-ml/TinyWord2-128k.

It solves the single-head attention collapse and character-stutter issues present in MicroWord-23k.

Benchmark Comparison

Metric Original MicroWord-23k DistilWord-23k (Ours)
Validation Loss 6.4105 3.0633
Validation PPL 608.18 21.40 (Lower = Better)

Generations from the model

Prompt Original MicroWord-23k DistilWord-23k
w wwwhwwwwgandwlwss wxing
z zzzzx's zred
el elel ely
app appco applval
Downloads last month
20
Safetensors
Model size
23.4k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support