Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Ran Le
leran1995
40
3
11
Follow
Yingzir's profile picture
vviiY's profile picture
xTimeCrystal's profile picture
14 followers
·
1 following
AI & ML interests
LLM Pretraining
Recent Activity
updated
a model
2 days ago
Nanbeige/Nanbeige4.2-3B
new
activity
2 days ago
Nanbeige/Nanbeige4.2-3B:
Amazing for its size, but spirals into questionable solutions once it fails something
new
activity
4 days ago
Nanbeige/Nanbeige4.2-3B:
Massive activation at layer 21, and three follow-ups on loop depth + LoopSplit
View all activity
Organizations
leran1995
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
Nanbeige/Nanbeige4.2-3B
2 days ago
Amazing for its size, but spirals into questionable solutions once it fails something
1
#21 opened 2 days ago by
Mk2Oracle
New activity in
Nanbeige/Nanbeige4.2-3B
4 days ago
Massive activation at layer 21, and three follow-ups on loop depth + LoopSplit
1
#20 opened 4 days ago by
an0nya
Suitability for 8 GB graphics cards
👀
4
3
#18 opened 4 days ago by
thermi6
Undocumented config params + two questions on loop depth and length-control RL
1
#16 opened 5 days ago by
an0nya
New activity in
Nanbeige/Nanbeige4.2-3B
5 days ago
Some questions about training
1
#14 opened 6 days ago by
mobinx
New activity in
Nanbeige/Nanbeige4.2-3B-Base
6 days ago
Request for Access to the Pre-training Dataset for Nanbeige/Nanbeige4.2-3B-Base (Open-Source Research)
👍
1
2
#1 opened 6 days ago by
arpitsh018
New activity in
Nanbeige/Nanbeige4.2-3B
6 days ago
the modal is tiny but kv cache exploding!
👍
6
2
#10 opened 8 days ago by
rosspanda0
New activity in
Nanbeige/Nanbeige4.2-3B
8 days ago
preserve_thinking
5
#9 opened 8 days ago by
owao
Please specify max context size in model's card.
1
#8 opened 8 days ago by
Reverger
Any plans for FP8?
2
#5 opened 8 days ago by
spanspek
Add community evaluation results for CLAW-EVAL, GPQA, HLE, HMMT_FEB_2026, SWE-BENCH_PRO, SWE-BENCH_VERIFIED, TERMINAL-BENCH-2.0
#7 opened 8 days ago by
nielsr
"Weimplify is asked:"
4
#4 opened 9 days ago by
owao
New activity in
Nanbeige/Nanbeige4.2-3B
9 days ago
Q on the model
4
#3 opened 9 days ago by
TomLucidor
gguf
➕
5
4
#1 opened 9 days ago by
Tapka
damn i was in the process of doing something very similar
1
#2 opened 9 days ago by
nraxl1
New activity in
Nanbeige/Nanbeige4.1-3B
about 1 month ago
Any plans for a larger scale up? (e.g., 7B - 12B version)
1
#46 opened about 1 month ago by
rpopreapovle
New activity in
Nanbeige/Nanbeige4.1-3B
5 months ago
Update README.md
#38 opened 5 months ago by
kerasakit
Overthinking Problem
➕
1
4
#27 opened 5 months ago by
JainilGosalia
Add syntax highlight in python markdown code snippets
#34 opened 5 months ago by
RahulSharma0
increase the max_new_tokens in the demo
#33 opened 5 months ago by
bitsnaps
Load more