picur-tokenizer / tokenizer.model

Commit History

a2C vocabulary: refit on the cleaned, budget-proportional corpus
a858237
verified

lazos commited on

a2B: reproducible seed, public benchmarks, benchmark.py
4079ffe
verified

lazos commited on

a2u: 28,000-stem seed, measured on text withheld from the fit
a34a31c
verified

lazos commited on

picur-tokenizer a2x: 49,152-piece Hungarian Unigram vocabulary
a7cf8a7
verified

lazos commited on