👋 Open to Work
C.J. Pitchford
databoyface
AI & ML interests
AGI, Education
Recent Activity
reacted to HannesVonEssen's post with 😎 8 days ago
📣 HF Viewer now supports Hugging Face login! 🤗
⚡ Generate visualizations for all your models at once!
😎 Feature your models in the community showcase and set up your own profile/org collection page on hfviewer!
✍️ Write your own model articles on hfviewer in the same novel interactive style as our "Gemma 4 family" article - with linking between the graph nodes and article text!
📊 We are also now rolling out support for tensor shapes, FLOPs and param counts per layer! :)
Thanks for all your positive feedback and suggestions! ❤️
Try the logged in experience here: https://hfviewer.com/
Here are some really cool articles already written by our users:
📑 LFM2.5-Audio: edge-first speech inference by Anna Piunova at Liquid AI
https://hfviewer.com/LiquidAI/LFM2.5-Audio-1.5B
📑 DeepSeek V4 mHC Explained by Shakti Wadekar
https://hfviewer.com/deepseek-ai/DeepSeek-V4-Pro
📑 Borealis - open recipe for training Audio LLMs by Alex Wortega
https://hfviewer.com/Vikhrmodels/Borealis-5b-it reacted to HannesVonEssen's post with 👍 8 days ago
📣 Understand any Hugging Face model!
💬 You can now hover a node to get a nice animated explanation of how the operation works and which paper introduced it! https://hfviewer.com
🎞️ Also, embedded HF Viewer graphs now default to being animated! Create and customize your animation at
https://hfviewer.com/model-card-embed
👥 Finally, feel free to join our discord to influence the future direction of HF Viewer! https://discord.gg/a5eEmtTTPV reacted to SeaWolf-AI's post with 👍 8 days ago
A small gift for anyone building or studying foundation models.
Most "open" models hand you the weights and stop there. With Aether-7B-5Attn we wanted to hand over the whole thing — so you can actually learn from it, reproduce it, and build on it: the data recipe, the training code, every hyperparameter, the complete logs, and the intermediate checkpoints. All Apache-2.0, reproducible byte-for-byte.
What you can do with it:
🔁 Rebuild it from scratch, or fork the recipe for your own model
🔬 Study a real heterogeneous-attention MoE — 49 layers place 5 attention mechanisms on a 7×7 Latin square, arranged as a clean, attributable ablation
📈 Trace training dynamics across the released checkpoints (110k / 115k / 162k)
It's a modest 6.59B model, and an honest one — the limitations (no KV-cache in this build, small scale) are written right in the card. We're not claiming it's special. If any piece of it saves you time or teaches you something, that's exactly what we hoped for. 🤗
📖 Full write-up →
[blog] · https://huggingface.co/blog/FINAL-Bench/opensource-llm
📦 5 Attention Base · https://huggingface.co/FINAL-Bench/Aether-7B-5Attn
🎯 5 Attention Instruct · https://huggingface.co/FINAL-Bench/Aether-7B-5Attn-it
🚀 5 Attention Live demo · https://huggingface.co/spaces/FINAL-Bench/Aether-Sovereign-AI
📦 7 Attention Base · https://huggingface.co/FINAL-Bench/Aether-7B-7Attn-base
📦 11 Attention Base · https://huggingface.co/FINAL-Bench/Aether-6B-11Attn-base
🧬 Collection · https://huggingface.co/collections/FINAL-Bench/aether-foundation-model
#opensource #LLM #MoE #reproducibility #Apache2Organizations
None yet