AI & ML interests

Exploring artificial intelligence together. Open source

Recent Activity

wop  updated a dataset about 15 hours ago
bench-labs/slop-classification
wop  updated a model about 15 hours ago
bench-labs/objectmodel-v1
TobiasLogic  updated a collection about 20 hours ago
Roadmap
View all activity

wop 
updated a Space 4 days ago
wop 
posted an update 4 days ago
view post
Post
2145
bench-labs/GCTokenizer-v1 , a multilingual tokenizer which does not require a training corpus

bench-labs
developed **GCTokenizer-v1**, which is a multi-lingual tokenizer
Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL"
It utilizes an encoding scheme which allows it to handle characters in any language around the world

General (multi lingual)
Consensus (from multiple model tokenizers consensus)
Tokenizer

We included an implementation script too,
built like BPE- it can encode arbitrary text, most of the time, efficiently
wop 
posted an update 5 days ago