·
AI & ML interests
Working on small sota models
Recent Activity
reacted to ucr-max's post with 👍 about 16 hours ago Introducing Limen0.2B
We are releasing Limen0.2B, a 222.5M-parameter base language model developed as a research platform for efficient pretraining and superword tokenization at smaller scales.
Limen0.2B was trained from scratch on 50B tokens and uses a compact 16K BoundlessBPE vocabulary. The project explores whether SuperBPE-style tokenization can remain effective in a substantially smaller model and vocabulary regime than those examined in earlier large-scale experiments.
The model also combines a deep-and-narrow transformer design with Exclusive Self-Attention, grouped-query attention, and tied embeddings. Its compact vocabulary reduces the embedding footprint and leaves a larger share of the parameter budget available to the transformer layers.
Despite its relatively modest training budget, Limen0.2B achieves competitive results for its scale across the reported language understanding, commonsense reasoning, and grammatical evaluation tasks. Comparisons with other compact models are provided as context rather than strict rankings, as their training data, token budgets, architectures, and evaluation settings differ.
The release includes the model weights, implementation, training configuration, checkpoint progression, and evaluation results, all under Apache 2.0.
https://huggingface.co/UniversalComputingResearch/Limen0.2B
Technical feedback, independent evaluations, and further experiments with the model and tokenizer are welcome.
View all activity Organizations
Text Generation
• 0.5B • Updated • 11
appvoid/graphite-001-large-Q8_0-GGUF
Text Generation
• 0.4B • Updated • 9
appvoid/graphite-001-large
Text Generation
• 0.4B • Updated • 11
appvoid/Falcon-H1-Tiny-Multilingual-100M-Instruct-Q8_0-GGUF
0.1B • Updated • 10
appvoid/Falcon-H1-Tiny-Multilingual-100M-Base-Q8_0-GGUF
0.1B • Updated • 12
appvoid/Falcon-H1-Tiny-90M-Instruct-pre-DPO-Q8_0-GGUF
91.1M • Updated • 5
appvoid/Falcon-H1-Tiny-90M-Instruct-Curriculum-pre-DPO-Q8_0-GGUF
91.1M • Updated • 9
appvoid/Falcon-H1-Tiny-90M-Instruct-Curriculum-Q8_0-GGUF
91.1M • Updated • 5
appvoid/Falcon-H1-Tiny-90M-Base-Q8_0-GGUF
91.1M • Updated • 6
0.1B • Updated • 10
0.3B • Updated • 11
appvoid/granite-4.0-350m-minimal-mix-sft-lora
Updated
appvoid/granite-4.0-350m-minimal-mix-sft-merged
Text Generation
• 0.4B • Updated • 10
appvoid/functiongemma-hermes-3k-ft-Q8_0-GGUF
Text Generation
• 0.3B • Updated • 7
• 1
Text Generation
• 0.6B • Updated • 4
Text Generation
• 0.6B • Updated • 3
Text Generation
• 0.6B • Updated • 4
Text Generation
• 0.6B • Updated • 3
appvoid/Qwen3-0.6B-notetaker-Q8_0-GGUF
0.6B • Updated • 8
appvoid/qwen3-0.6b-refiner-codeql-self-nothink-full-Q8_0-GGUF
0.6B • Updated • 4
appvoid/RAGU-lm-Q8_0-GGUF
0.6B • Updated • 7
appvoid/Qwen3-0.6B-Shadow-FT-BAAI-2k-Q8_0-GGUF
0.6B • Updated • 2
appvoid/Dripper-Q8_0-GGUF
Text Generation
• 0.8B • Updated • 3
appvoid/event-attribute-extractor-Q8_0-GGUF
0.6B • Updated • 3
appvoid/Qwen3-0.6B-Books-Intent-Q8_0-GGUF
0.6B • Updated • 11
appvoid/syngen-reasoning-0.6b-Q8_0-GGUF
Text Generation
• 0.6B • Updated • 8
appvoid/Qwen3-0.6B-m3-mcqa-reason-besthyper-fixed-Q8_0-GGUF
0.6B • Updated • 8
appvoid/qwen3_FineTome-100k-Q8_0-GGUF
0.6B • Updated • 7