Tardelli Stekel
AI & ML interests
Recent Activity
Organizations
Thanks for these datasets, and a question about the license
Thanks, this was worth digging into, and it changed the data.
The half stalls are real, and they're the same events as the dataset stalls. On those days the dataset total goes to zero, with every dataset frozen. On the model side only part of the counters freeze, and it splits by repo age: about 98% of models created before 2022 keep counting (BERT, CLIP, all-MiniLM-L6-v2), while almost no repo from mid-2025 on does. The old giants are what hold the total at half. 20 of the 21 such days are Wednesdays, and snapshot spacing is 24.0 h on all of them, so it's not timing.
08-13 is a counter catching up. On 08-12, 80% of models with 2k+ daily downloads were frozen at exactly zero (bge-small-en-v1.5, bge-m3, chronos-2, MiniMax-H3), while all-MiniLM-L6-v2 kept counting at twice its usual rate. On 08-13, 94% of them were above 1.8x their own median, and the top 20 explain only 30% of the excess. Frozen models do 2โ4x the next day, so these downloads are delayed, not lost.
Your pair rule catches most of them but not all. The share of frozen counters also catches stalls that barely dip the total (2025-12-17, 2026-05-12, 2026-09-02), and misses ones where counters crawled instead of stopping. 06-25 and 08-01 have nothing frozen and normal datasets, but every model sits at about 0.5x and then 1.6โ1.9x, so they read as a shifted count too. A day is now a stall if any of these holds: total under 0.3x of the median, at least 25% of models with 60K+ monthly downloads frozen, or your pair test (under 0.7x, made up by a neighbour to about 2x). Datasets use a 90% frozen bar, since they freeze all at once.
And 2025-05-21 wasn't a spike. On 05-19 and 05-20, dl_all went down for 16โ19% of models and came back days later. We clipped the drop but counted the rebound: 440M phantom downloads. The same happened from 2026-06-13, another 537M. Snapshots during a rollback are now set aside, and downloads are measured across them.
It's live. The model side has 27 windows and skip_days went from 20 to 59 (datasets 57 to 65). The Hub total dropped 0.99B to 43.48B, all of it from the two rollbacks, and days outside 0.7โ1.3x of the median went from 70 to 40. 08-01/02 is the one pair that still slips through, right at the edge of the sum band.
Well done
Thanks again, this was a good catch. I checked fetch times first: each day is the last hub-stats commit of that UTC day, so I recovered the timestamps from the commit history. Around these days the snapshots are still about 24h apart (about 13:10 UTC), only 2025-03-04 comes 10h after the previous one, and the files differ every day (new models keep landing).
So it's not two snapshots close together: the Hub's download counters stood still that day and caught up over the next day or two, sometimes on the day before. Weighting by elapsed hours wouldn't fix it.
What's in now: a day below 30% of its 15-day median is a stall. It's grouped with the catch-up days around it (above 1.3ร the median), and the window's total is spread evenly across its days, tag by tag. That gives 10 windows, your 7 days plus partial ones on 2025-05-03/04, 2025-09-10, 2026-04-01, 2026-06-12/13 and 2026-09-09.
Sums are unchanged (44.376B, rounding aside), and no day is left under 30% of its median.
Per-model series skip the snapshots inside each window (listed as skip_days in meta.json), so model, family and author charts spread those days the same way. New stalls get settled automatically once their catch-up arrives.
Dataset commit: https://huggingface.co/datasets/modelpulse/model-pulse-data/commit/aaf4ef4c417ccad18d998e496772fba205337049
A model page tells you its downloads for the last 30 days, but not how it got there. Model Pulse rebuilds the day-by-day history for 1.6M models, back to July 2024, from the daily snapshots in @cfahlgren1 's hub-stats dataset.
A few things the data shows:
- The Hub now serves ~100M model downloads a day, up from ~62M a year ago
- Vision-language models grew about 7x in that year, to 10.7M downloads a day
- 37% of Qwen3-8B's monthly downloads go to its 6,283 quantizations, fine-tunes and merges, not to the original
What you can do with it:
- Open any model, e.g. tardellirs/model-pulse
- Compare up to five models on one chart
- Browse weekly rankings: fastest growing, new this month, biggest families, top organizations
- Add a live badge to your model card (monthly downloads, sparkline and weekly trend)
The data is open too: modelpulse/model-pulse-data
I'd love feedback, especially from model authors: what would you want to see about your own models?
Thanks a lot, you're right, and your numbers check out exactly: 17 gaps, 92 missing days, 6.48B downloads (14.6%) missing from row sums. I went with your first option. hub_series.parquet now has one row per calendar day, with each gap's per-day average written to every day it covers, so sums are exact and the charts get a continuous series. The file in modelpulse/model-pulse-data is already updated (583 consecutive days, total 44.4B), and the daily job writes gaps the same way from now on. The report's headline numbers came from month-end cumulative totals, so they're unchanged. Really appreciate you checking it against series/, that's exactly why the data is open.
A model page tells you its downloads for the last 30 days, but not how it got there. Model Pulse rebuilds the day-by-day history for 1.6M models, back to July 2024, from the daily snapshots in @cfahlgren1 's hub-stats dataset.
A few things the data shows:
- The Hub now serves ~100M model downloads a day, up from ~62M a year ago
- Vision-language models grew about 7x in that year, to 10.7M downloads a day
- 37% of Qwen3-8B's monthly downloads go to its 6,283 quantizations, fine-tunes and merges, not to the original
What you can do with it:
- Open any model, e.g. tardellirs/model-pulse
- Compare up to five models on one chart
- Browse weekly rankings: fastest growing, new this month, biggest families, top organizations
- Add a live badge to your model card (monthly downloads, sparkline and weekly trend)
The data is open too: modelpulse/model-pulse-data
I'd love feedback, especially from model authors: what would you want to see about your own models?