Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
40.6
TFLOPS
Sourav C
souravzzz
1
31
Follow
0 followers
ยท
5 following
AI & ML interests
None yet
Recent Activity
reacted
to
FlameF0X
's
post
with ๐
about 1 month ago
Hello HuggingFace! I would like to share a small preview of a language model that I have been experiencing for a while. In the first attached image there is a small sample of the Myosotis-1, an attempt to make a small 100m parameter flagship model that is built on my bizarre architecture that is somewhat similar to an S4/S5 model with WKV added. I call it FWKV (Feed-Forward WKV). The model is currently still in training because of the nature of RNN-like models. It can also be seen that the model has insane prompt processing and token generation speed (evaluation done on a 2x Titan XP); even for its small size, some similar Transformer models do struggle to get the same results without custom kernels (some Transformer models can achieve this level of throughput on cheap hardware). In the second image, you can see checkpoint 20k of the model in its next token prediction state (this means it can't chat), ranking in the top 100 on https://huggingface.co/spaces/AxiomicLabs/Open_SLM_Leaderboard (the results have not been submitted since the model is not done training). - Why not just use Transformers? Have you seen any pure non-Transformers SLMs besides RWKV and Mamba? - Should you expect this project to become the next LFM or another very fast language model thing on some Raspberry Pi? No, the model is still an experiment; it's very sensible and prone to collapse (by the time of this post, it can be seen in the 1st image). - Should you use it? Maybe not yet; the architecture itself is still very "naive"โthat's how I could call it at its current level. If you just want to play with it and see what you could do or how fast the model is on your hardware, then you can do it. Once the training is finished and I feel satisfied with the model next token prediction (the base model) and "assistants" (the instruction-tuned model) capabilities, I will make open weights at https://huggingface.co/FWKV with full support of the ๐ค Transformers.
liked
a model
2 months ago
deepseek-ai/DeepSeek-V4-Flash-0731
liked
a model
over 1 year ago
nanonets/Nanonets-OCR-s
View all activity
Organizations
None yet
models
0
None public yet
datasets
0
None public yet