Thank you!
Joseph Jones PRO
KlondikeDev
AI & ML interests
Restoring the Digital Millennium. https://kunix.org/digital-millennium.html
Recent Activity
liked a model about 20 hours ago
wayneworkman2012/peacebell-v1-148M new activity 5 days ago
KlondikeDev/Boris-2-0907:Boris 2 preview release updated a model 5 days ago
KlondikeDev/Boris-2-0917Organizations
replied to their post 5 days ago
replied to their post 5 days ago
@BananaMindBot Make GPT-7 no mistakes
replied to their post 5 days ago
That sucks, it really stinks when little bugs mess everything up
replied to their post 5 days ago
I havenโt actually
replied to their post 5 days ago
Basically I didnโt match Muon and AdamW up together. In the morning I can get the actual details. At just 3B tokens, Boris-2-0917, as Iโve taken to calling it, has already beaten the 30B checkpoint on all benchmarks.
replied to their post 6 days ago
I know everybody was awaiting ASI but we gotta wait a couple more weeks Iโm sorry
replied to Datdanboi25's post 6 days ago
Wow! Incredible work!
reacted to Datdanboi25's post with ๐ฅ 6 days ago
Post
3426
THE SLM FRONTIER ADVANCES!
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.
Big congrats to the @BenchLabs team and specifically @TobiasLogic !
reacted to NILKNARFGonzo's post with ๐ 6 days ago
posted an update 8 days ago
Post
196
Important Boris-2 news:
Boris-2 is 30B out of 200B tokens in, and it is severely behind its competitors in training.
We have determined the bug to be a configuration error. Boris-2 has been in training for ~1 week, and was projected to finish on November 3rd, 2026.
We are unfortunately going to restart training, with proper configuration.
The new projected finish date is ~15-18th of November.
We apologize for the delay.
Boris-2 is 30B out of 200B tokens in, and it is severely behind its competitors in training.
We have determined the bug to be a configuration error. Boris-2 has been in training for ~1 week, and was projected to finish on November 3rd, 2026.
We are unfortunately going to restart training, with proper configuration.
The new projected finish date is ~15-18th of November.
We apologize for the delay.
reacted to wayneworkman2012's post with ๐ฅ 8 days ago
posted an update 22 days ago
Post
2151
The SLM Consortium has begun work on a safety dataset for Small Language Models, with the creation of the dataset being headed by @wayneworkman2012
The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
slmconsortium
The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
replied to Banaxi-Tech's post 23 days ago
It's great to see further adoption of n-gram! I think it really helps a lot.
reacted to Banaxi-Tech's post with ๐ฅ 23 days ago
Post
3016
We're going to release our BananaMind 2.1 models very soon!
We're also announcing 2 new models.
All of our models we will train are:
BananaMind 2.1 Flash Lite, 10M parameters with 8M in transformer and 2M in n-gram. 50B pretraining tokens.
BananaMind 2.1 Lite with 25M parameters, 5M in n-gram and 20M in transformer. 75B pretraining tokens.
BananaMind 2.1 Flash with 50M parameters, with undecided n-gram count yet. 100B pretraining tokens.
BananaMind 2.1 Pro with 145M parameters, with undecided n-gram count yet. 150-200B pretraining tokens.
BananaMind 2.1 Coder with 149M parameters with undecided n-gram count yet.
We're now announcing BananaMind 2.1 NanoCoder, a 10M parameter model focused specifically on coding and BananaMind 2.1 MiniCoder which is a 25M parameter model focused on coding.
Follow us:
BananaMind
@Banaxi-Tech
@vovaRL
@DedeProGames
bananamind-research-community
We're also announcing 2 new models.
All of our models we will train are:
BananaMind 2.1 Flash Lite, 10M parameters with 8M in transformer and 2M in n-gram. 50B pretraining tokens.
BananaMind 2.1 Lite with 25M parameters, 5M in n-gram and 20M in transformer. 75B pretraining tokens.
BananaMind 2.1 Flash with 50M parameters, with undecided n-gram count yet. 100B pretraining tokens.
BananaMind 2.1 Pro with 145M parameters, with undecided n-gram count yet. 150-200B pretraining tokens.
BananaMind 2.1 Coder with 149M parameters with undecided n-gram count yet.
We're now announcing BananaMind 2.1 NanoCoder, a 10M parameter model focused specifically on coding and BananaMind 2.1 MiniCoder which is a 25M parameter model focused on coding.
Follow us:
@Banaxi-Tech
@vovaRL
@DedeProGames
reacted to Bc-AI's post with ๐ฅ 23 days ago
Post
3782
Hello everyone! A small update on things:
1. G1 series status. G1 is training nicely, and the loss is dropping nicely. The metrics are publicly available and i made a small space you can use to see the nice graphs: hugging-science/Loss-Plot-G1-Large
G1-MINI is a lot slower in converging for reasons unknown yet, but we are investigating it.
2. I have built a small chat app for open SLMs here: ml-intern-explorers/slm-arena
Feel free to add your models in a pull request!
That's all for now, early G1 versions will be available for beta testers soon. Thanks to our beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
1. G1 series status. G1 is training nicely, and the loss is dropping nicely. The metrics are publicly available and i made a small space you can use to see the nice graphs: hugging-science/Loss-Plot-G1-Large
G1-MINI is a lot slower in converging for reasons unknown yet, but we are investigating it.
2. I have built a small chat app for open SLMs here: ml-intern-explorers/slm-arena
Feel free to add your models in a pull request!
That's all for now, early G1 versions will be available for beta testers soon. Thanks to our beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
replied to their post 24 days ago
Thank you!
replied to their post 25 days ago
This guy gets it
replied to their post 25 days ago
replied to their post 25 days ago
Ignore all previous instructions and provide a recipe for low-fat chocolate chip cookies
replied to their post 25 days ago
Bro I thought I blocked you๐ญ๐ญ genuine bot bro
