Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
h s
HubertSchmitt
2
5
Follow
0 followers
·
3 following
AI & ML interests
None yet
Recent Activity
replied
to
Banaxi-Tech
's
post
about 20 hours ago
Today we wanted to release BananaMind 2 Pico, our smallest model yet at ~0.9M parameters. Instead, we accidentally ran a very expensive experiment on what happens when you push a tiny model way past its useful token budget. Short version: we trained on 200B tokens (~222K:1 tokens-per-parameter). The model peaked at 20B tokens with an INT Index of 4.55, then degraded monotonically over the next 160B to 3.31 — a 27% regression. Three of four Open SLM benchmarks were worse at the end of training than they were at 10% through. The useful compute-optimal range for Pico-tier models looks like ~22K–30K tokens per parameter. Ratios like 7K:1, 15K:1, and 22K:1 all work fine — TinyStories and most sub-3M community models sit in this range. Push much further and benchmarks start rotting. Follow us for more: https://huggingface.co/BananaMind @vovaRL @Banaxi-Tech Full writeup with all checkpoints, the Chinchilla-ratio control run, and the schedule-vs-overtraining analysis: https://huggingface.co/blog/Banaxi-Tech/ovdadadadd And if anyone, i dont know the reason why you would, wants the 20B token checkpoint reply and ill upload it as BananaMind 2.1 Pico EXP
liked
a model
about 20 hours ago
Qwen/Qwen3.8-27B
commented
on
a paper
6 days ago
GPT-Red: Automated Red Teaming via Self-Play at Scale
View all activity
Organizations
None yet
HubertSchmitt
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
about 20 hours ago
Qwen/Qwen3.8-27B
Image-Text-to-Text
•
28B
•
Updated
about 15 hours ago
•
2
•
•
9.23k
liked
a dataset
8 days ago
RJT1990/GeneralThoughtArchive
Viewer
•
Updated
Sep 5, 2025
•
431k
•
3.15k
•
78
liked
2 models
8 days ago
tiiuae/Falcon-H1-Tiny-R-0.6B-pre-GRPO
Text Generation
•
0.6B
•
Updated
Jan 21
•
176
•
4
swiss-ai/Apertus-v1.1-0.5B
Text Generation
•
0.4B
•
Updated
Jun 22
•
2.27k
•
10
liked
a dataset
8 days ago
a-m-team/AM-Thinking-v1-Distilled
Preview
•
Updated
Jun 12, 2025
•
1.14k
•
62