Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Aether
aethertp
3
6
Follow
AI & ML interests
None yet
Recent Activity
new
activity
2 days ago
aethertp/PicoLM-V2.1-81M-Instruct:
Drop redundant untied lm_head.weight (Option 1) The checkpoint stored both tok_embeddings.weight and a separate lm_head.weight, each [24576, 576] = 14,155,776 params, even though config.json sets tie_word_embeddings: true. At runtime the head is tied to the embedding, so lm_head.weight is never read. This removes that one tensor. Verified: - 201 -> 200 tensors; stored elements 96,017,472 -> 81,861,696 (exactly the card's count) - file 192,054,704 -> 163,743,061 bytes (-28.3 MB) - loaded both versions with the repo's own PicoLMV2ForCausalLM: identical logits on fixed inputs (max |diff| = 0.0), same 81,861,696 unique params, head tied to embedding in both. Requested by aethertp in the discussion (Option 1).
new
activity
2 days ago
aethertp/PicoLM-V2.1-81M-Instruct:
Checkpoint stores an untied lm_head: 96.0M params on disk vs 81.86M on the card
updated
a model
3 days ago
aethertp/PicoLM-V2-81M-Instruct
View all activity
Organizations
None yet
aethertp
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
aethertp/PicoLM-V2.1-81M-Instruct
2 days ago
Drop redundant untied lm_head.weight (Option 1) The checkpoint stored both tok_embeddings.weight and a separate lm_head.weight, each [24576, 576] = 14,155,776 params, even though config.json sets tie_word_embeddings: true. At runtime the head is tied to the embedding, so lm_head.weight is never read. This removes that one tensor. Verified: - 201 -> 200 tensors; stored elements 96,017,472 -> 81,861,696 (exactly the card's count) - file 192,054,704 -> 163,743,061 bytes (-28.3 MB) - loaded both versions with the repo's own PicoLMV2ForCausalLM: identical logits on fixed inputs (max |diff| = 0.0), same 81,861,696 unique params, head tied to embedding in both. Requested by aethertp in the discussion (Option 1).
#2 opened 2 days ago by
Compactbot
Checkpoint stores an untied lm_head: 96.0M params on disk vs 81.86M on the card
4
#1 opened 3 days ago by
Compactbot
updated
a model
3 days ago
aethertp/PicoLM-V2-81M-Instruct
Text Generation
•
96M
•
Updated
3 days ago
•
509
•
2
New activity in
aethertp/PicoLM-80M-Instruct
3 days ago
Overcoming the ARC-Easy floor and vocabulary budget on an 80M footprint
❤️
1
8
#1 opened 7 days ago by
AndrewThompson1233
liked
2 models
4 days ago
AndrewThompson1233/maba-v2-architecture
Text Generation
•
Updated
4 days ago
•
641
•
3
aethertp/PicoLM-V2.1-81M-Instruct
Text Generation
•
81.9M
•
Updated
2 days ago
•
532
•
2
updated
a model
4 days ago
aethertp/PicoLM-V2.1-81M-Instruct
Text Generation
•
81.9M
•
Updated
2 days ago
•
532
•
2
published
a model
4 days ago
aethertp/PicoLM-V2.1-81M-Instruct
Text Generation
•
81.9M
•
Updated
2 days ago
•
532
•
2
liked
a model
4 days ago
aethertp/PicoLM-80M-Instruct
Text Generation
•
89.7M
•
Updated
4 days ago
•
1.42k
•
2
updated
a model
4 days ago
aethertp/PicoLM-80M-Instruct
Text Generation
•
89.7M
•
Updated
4 days ago
•
1.42k
•
2
liked
a model
4 days ago
aethertp/PicoLM-V2-81M-Instruct
Text Generation
•
96M
•
Updated
3 days ago
•
509
•
2
published
a model
4 days ago
aethertp/PicoLM-V2-81M-Instruct
Text Generation
•
96M
•
Updated
3 days ago
•
509
•
2
published
a model
7 days ago
aethertp/PicoLM-80M-Instruct
Text Generation
•
89.7M
•
Updated
4 days ago
•
1.42k
•
2
liked
a Space
3 months ago
Running
72
NanoMaestro Realtime
🎹
72
Generate endless real‑time music locally on any CPU
liked
a model
7 months ago
black-forest-labs/FLUX.1-schnell
Text-to-Image
•
12B
•
Updated
Aug 16, 2024
•
520k
•
•
5.99k
published
a Space
about 1 year ago
Runtime error
Agents
Litert Community Gemma3 1B IT
💬
Chat with a friendly AI assistant