AI & ML interests

Breaking the opacity of language models for legal professionals ๐Ÿ“– Join us by smashing the button at top right ๐Ÿค—

Recent Activity

Nymboย 
posted an update 2 days ago
view post
Post
110
FYI All my spaces are confirmed working again and actively maintained now. If you tried one a while back and it was broken, it should work now. I'm also committed to keeping up with issues and PRs going forward. If something's still busted, open an issue on the Space and I'll actually see it :)
lianghsunย 
posted an update about 1 month ago
view post
Post
315
๐Ÿ‡น๐Ÿ‡ผ Releasing https://huggingface.co/lianghsun/tw-tokenizer-v1 โ€” a tokenizer trained from scratch for Traditional Chinese (Taiwan).

**46% better Chinese compression than Qwen3.8-27B with 81% of its vocab (201K vs 248K), and English essentially untouched (4.657 vs 4.674 chars/token).**

The gain isn't from the regex โ€” it's the corpus. Qwen carries **27,364 Simplified-only multi-char tokens**, 11% of its vocab, dead weight for Traditional Chinese. Train on pure Traditional and that waste never appears.

Recent work is skeptical that compression predicts quality (Lotz et al. 2025 measured ฯ = โˆ’0.59), so we validated two levels deeper:

**Segmentation** โ€” boundary hit rate against jieba: **85.6%** vs Qwen's 77.8%. Single-character tokens: **17.6%** vs 41.7%.

ๅฐˆๆฅญ็ด ้คŠใ€็‰น่ณชๆˆ–็ถ“ๅ…ฌๅ‘ŠๅฏฉๆŸฅๅ„ชๅ‹
  ours: ['ๅฐˆๆฅญ็ด ้คŠ', 'ใ€', '็‰น่ณช', 'ๆˆ–็ถ“', 'ๅ…ฌๅ‘Š', 'ๅฏฉๆŸฅ', 'ๅ„ชๅ‹']
  Qwen: ['ๅฐˆๆฅญ', '็ด ', '้คŠ', ...]     โ† ใ€Œ็ด ้คŠใ€split mid-word


**Downstream** โ€” trained a 270M model from scratch with each tokenizer, compared bits-per-character (the only metric fair across tokenizers). At equal compute: **4.434 vs 4.591**, a 3.4% win โ€” with 13% fewer parameters. Same token budget means our model saw 440M characters vs 308M: **43% more data for the same compute**.

Also: 6-char cap on pure-CJK tokens (long tokens obscure orthographic info โ€” Haslett, CL 2025), NFC not NFKC, 1,024 reserved tokens.

Known limits (weak Tรขi-lรด support, small-scale downstream validation, vocab sweep hadn't flattened) are in the card.

๐Ÿ‘‰ https://huggingface.co/lianghsun/tw-tokenizer-v1
Nymboย 
posted an update about 2 months ago
view post
Post
2410
Anthropic gave me six months of Claude Max 20x through the Claude for Open Source program, granted based on my Hugging Face work. Thank you
Anthropic
for supporting open source.

So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.

https://github.com/Nymbo/Markdown-Minimap โ€” issues and PRs welcome.
Nymboย 
posted an update 2 months ago
view post
Post
6103
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.

CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.

See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2

Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
  • 6 replies
ยท
eienmojikiย 
posted an update 3 months ago
Tonicย 
posted an update 5 months ago
view post
Post
3443
๐Ÿ™‹๐Ÿปโ€โ™‚๏ธ Hey there folks ,

Turns out : if we predict ๐ŸŒ earth we can save a lot of time looking for interesting things and less time looking at things that we expect to see.

Sentinel-2 imagery ๐Ÿ›ฐ๏ธbasically takes a long time to download towards earth. so our "near real time" systems are quite far from that in practical terms.

meanwhile , if we "predict" what we will see , based on what we do see , we can send down much less data in a timely way , and prioritize ๐Ÿ“กearth-bound response .

I'm talking about illegal fishing , logging , mining or building in nature reserves , the more of that we predict early the more we're able to stop it on time.

At least that's the concept !

check out the blog : https://huggingface.co/blog/Tonic/save-patagonia-by-predicting-earth


- Collection: https://huggingface.co/collections/NuTonic/earth-observation-with-temporal-and-general-understanding
- Code: https://github.com/Josephrp/Nutonic
- Dataset: NuTonic/sat-vl-sft-training-ready-v1
- Model: NuTonic/lspace
- Training: NuTonic/lspace-trackio
- Evals: NuTonic/Patagonia_Eval
  • 2 replies
ยท
Tonicย 
posted an update 5 months ago
view post
Post
4506
๐Ÿ™‹๐Ÿปโ€โ™‚๏ธ Hey there folks,

since everyone liked my previous announcement post ( https://huggingface.co/posts/Tonic/338509028435394 ) so much , i'm back with more high quality proceedural datasets in the Geospacial domain for SFT training !

Check this one out :
NuTonic/sat-bbox-metadata-sft-v1

the goal is to be able to train vision models on multiple images for remote sensing analysis with one shot .

hope you like it ! ๐Ÿš€
  • 2 replies
ยท
Tonicย 
posted an update 6 months ago
view post
Post
3755
๐Ÿ™‹๐Ÿปโ€โ™‚๏ธ Hey there folks ,

I'm sharing huggingface's largest dataset of annotated statelite images today.

check it out here : NuTonic/sat-image-boundingbox-sft-full

I hope you like it , the idea is to be able to use this with small vision models ๐Ÿš€
umarbutlerย 
posted an update 7 months ago
view post
Post
5006
Isaacus, the AI research company building legal superintelligence, is hiring!

We're looking for passionate engineers who love to build and tinker and want to have an impact on the world. Specifically, we're hiring:
โ€ข ML engineers (Australia).
โ€ข Data engineers (Australia).
โ€ข Full-stack engineers (Australia).
โ€ข DevRel engineers (Australia, San Francisco, and London).
โ€ข DevOps engineers (Australia, San Francisco, and London).

If you'd like to be a founding employee at one of the few VC-backed LLM research labs in the world, receive generous equity compensation, and work alongside other highly motivated, highly skilled engineers, get in touch: https://isaacus.com/careers