Dot Labs

company
Activity Feed

AI & ML interests

Multimodal, Byte-Level research

Recent Activity

appvoid  updated a Space 3 days ago
dotlabs/README
appvoid  updated a dataset 3 days ago
dotlabs/rewrite
appvoid  published a dataset 13 days ago
dotlabs/rewrite
View all activity

appvoid 
updated a Space 3 days ago
appvoid 
posted an update 3 days ago
view post
Post
134
Reality has not limits if you know how to shape it, one step at a time.
  • 1 reply
·
appvoid 
posted an update 4 days ago
view post
Post
68
How old were you when you discovered good AI agents favor the Strategy Design Pattern over any other?
  • 1 reply
·
appvoid 
posted an update 15 days ago
view post
Post
127
LLMS for edge devices?
What if we go the reliable route instead of the speed route?
What if we make it run on sbcs with few megabytes available?

That's the idea for the next model.

Keep in tune.
appvoid 
posted an update 18 days ago
view post
Post
137
We trained a 10.9M byte-level recurrent Transformer on L3 and L6. (Loop 3 and Loop 6)

Yet L4/L5 improved too, L8 held up, and the L3→L6 gain grew during training.

Same weights. More compute. Better predictions.

This is a new architecture for effective compute after several steps beyond original training!

We mixed and matched components like time and mhc into an ouro-like byte-level language model and the result is BET, a byte-level step-elastic transformer that can run computation steps without significant degradation.

One of the coolest parts of this training was discovering how Gradient Descent decided to use the first layer as what we would consider a scratchpad! Totally destroyed for the decoder but somehow makes total sense for the next layer!

I believe looped-transformers are the future of edge computing and this is a first step towards it.

Blogpost: https://medium.com/@appvoidofficial/byte-level-elasticity-182fe2ed1d2f

appvoid/bet-10m
appvoid 
posted an update 23 days ago
view post
Post
3900
We got gpt6 before gta6
  • 15 replies
·
appvoid 
published a Space 24 days ago
appvoid 
posted an update 25 days ago
view post
Post
126
Any thoughts on Nvidia acquiring this website?

I don't know what to feel about it. But would be great if huggingface gets something similar to Kaggle with free GPU hours (or even days) for training.
  • 3 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
3304
Love how the small lm community is getting identity over time:

- Channel-Mixing
- XSA
- Three-tower
- Digit aware
- Loops

No one is doing the same! That's so cool.
  • 36 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
2650
Nobody knows what is doing, when you train a model, you are experimenting to advance the frontier, so keep failing 🫵
  • 17 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
2603
I hope that after this OpenAI disaster on Plus users, more people start realizing why Open Weights were always the only way.
  • 5 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
1004
Byte-level state-space models. That sounded pretty scary for a scientist decades ago. Now we have:

1. Knowledge that deeper layers train smoothly.
2. Knowledge that Transformers work but is quadratic on sequence length.
3. Knowledge that SSMs work even better. Numerically unstable sometimes.
4. Speculative-decoding.
5. Open high-quality data.
6. Knowledge that KD works.

It slowly feels like is no longer a bad idea.
  • 3 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
167
It's 2026 and there are no instantaneous/fast vision language models for cpus yet. That's another free idea.
  • 6 replies
·
appvoid 
posted an update about 1 month ago
view post
Post
144
CaraArchive highlights a broader reality of putting data online: once something is publicly accessible, it becomes extremely difficult to guarantee that it will remain under your control.

If there is information or artwork that you absolutely do not want copied, archived, scraped, downloaded, or used by others, the safest option is still not to publish it publicly in the first place. That may sound obvious, but the internet was fundamentally designed to move and reproduce information, and there are countless ways to retrieve publicly accessible images:from ordinary browser tools and web scraping to automated or agentic systems.

That does not mean artists should simply accept every possible use of their work. Artists deserve meaningful control, attribution, compensation, and reasonable ways to express how their work may be used. But treating the technology and peopple using it itself as the enemy is unlikely to solve the underlying problem.

There probably isn't a technical solution that can make a publicly visible image simultaneously viewable by everyone and impossible to copy. The realistic goal should therefore be to create better norms, incentives, licensing systems, and tools around how that content is used.

Like, we can imagine a future where every artist gets his/her own credentials and some kind of fingerprint done just like blockchain works. But that requires substantial cooperation among organizations, companies and individuals.

Technology and art are not inherently opposing sides though.
appvoid 
posted an update about 2 months ago
view post
Post
176
If you lack ideas for a cool model, here's one.

Train a model from scratch on wikipedia with one twist: the tokenizer changes the actual token ids used on every sample fed. If somehow still learns English, you have made an astonishing discovery.

You would have answered the question: Can a model learn human languages from structure alone?
  • 2 replies
·
appvoid 
posted an update about 2 months ago
view post
Post
1972
Random corporate secret of tonight:

Try overfitting a tiny model on billions of high-quality datapoints: you can't. You can do 100 epochs and see the model still improving.

You're welcome.
  • 1 reply
·
appvoid 
posted an update about 2 months ago
view post
Post
865
- GLM 5.2
- Flux 3
- New Qwen model
- New small model leaderboards
- Lots of people finetuning smol models.
- Some even under 12 year olds clauders are here (was not on my bingo card this year)
- ChatGPT's Sol became a lot faster this week
- LFM2.5 2.6b
- Kimi K3 (though only a few will run it)
- New Ling 3.0 Tiny
- New video model that is making south park videos?
- Deepseek v4 flash being more honest than bigger models
- The new model from meta


Everything Everywhere All At Once
  • 1 reply
·
appvoid 
posted an update about 2 months ago
view post
Post
2391
I don't know if it was us or one of you guys or maybe all of us at once but lately we have seen a finetuning/pretraining explosion of models below 200m params and we can't be more happy about it keep coming tinkerers all of this is possible because of you!
  • 8 replies
·