Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
6.8
TFLOPS
Pablo
p-ferrando
1
23
Follow
wbf8d797ec1df44c67's profile picture
Gofar13's profile picture
2 followers
·
14 following
AI & ML interests
None yet
Recent Activity
reacted
to
pavle-scalably
's
post
with 🔥
12 days ago
14 days serving https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4 to production agents on 2x RTX 5090 (vLLM 0.27, TP=2, 262K context, FP8 KV): 28,097 requests, 860.6M prompt tokens, 82.6% prefix-cache hit rate, TTFT p50 0.61 s, 0 engine errors. The observation: prefix cache, not throughput, decides whether a 27B model keeps up with agents. Mean request is 30,100 tokens in, 983 out, because every turn resends the whole session. Two flags mattered most: --max-num-seqs 12 (queue p95 went 9.4 s to 233.6 s past that) and --watermark 0.08 (preemptions 29 to 2). And thinking off for tool loops: 917 tokens in 11.7 s vs 11,170 in 144 s, same answer. Full config and counters: scalably.io/blog/qwen3-8-27b-nvfp4-rtx-5090-production Next we are preparing an 8x B300 node in an EU data center for open-weight serving. Which models or workloads are underserved for you?
liked
a model
15 days ago
prism-ml/Ternary-Bonsai-2-27B-gguf
liked
a model
23 days ago
deepseek-ai/DeepSeek-V4.1-Flash
View all activity
Organizations
None yet
p-ferrando
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
15 days ago
prism-ml/Ternary-Bonsai-2-27B-gguf
Text Generation
•
27B
•
Updated
7 days ago
•
3.87M
•
2.36k
liked
a model
23 days ago
deepseek-ai/DeepSeek-V4.1-Flash
Image-Text-to-Text
•
763B
•
Updated
1 day ago
•
768k
•
•
4.02k
liked
2 models
about 2 months ago
Qwen/Qwen3.8-27B
Image-Text-to-Text
•
28B
•
Updated
Aug 14
•
6.93M
•
•
16.8k
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation
•
304B
•
Updated
Aug 1
•
4.49M
•
•
4.01k
liked
a model
2 months ago
Nanbeige/Nanbeige4.2-3B
Text Generation
•
4B
•
Updated
22 days ago
•
74.8k
•
807
liked
3 models
4 months ago
sapientinc/HRM-Text-1B
Text Generation
•
1B
•
Updated
29 days ago
•
19.8k
•
824
LiquidAI/LFM2.5-8B-A1B
Text Generation
•
8B
•
Updated
Aug 24
•
33.1k
•
779
FINAL-Bench/Darwin-28B-Opus
Text Generation
•
28B
•
Updated
Jul 23
•
92
•
•
43
liked
a model
5 months ago
openbmb/MiniCPM-V-4.6
Image-Text-to-Text
•
1B
•
Updated
Aug 17
•
305k
•
1.23k
liked
2 models
6 months ago
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text
•
36B
•
Updated
Apr 24
•
3.24M
•
•
2.91k
FINAL-Bench/Darwin-31B-Opus
Text Generation
•
31B
•
Updated
Aug 11
•
103
•
•
65
liked
a model
7 months ago
mistralai/Mistral-Small-4-119B-2603
119B
•
Updated
Jul 15
•
54.7k
•
430
liked
8 models
9 months ago
baichuan-inc/Baichuan-M3-235B
Text Generation
•
235B
•
Updated
Feb 9
•
470
•
104
mistralai/Mistral-Large-3-675B-Instruct-2512
Updated
Jul 15
•
1.71k
•
252
mistralai/Ministral-3-14B-Instruct-2512
14B
•
Updated
Jul 15
•
269k
•
329
deepseek-ai/DeepSeek-V3.2-Speciale
Text Generation
•
685B
•
Updated
Dec 1, 2025
•
3.67k
•
726
tencent/Youtu-LLM-2B
Text Generation
•
2B
•
Updated
Feb 24
•
11.3k
•
231
LiquidAI/LFM2.5-1.2B-Instruct
Text Generation
•
1B
•
Updated
Aug 24
•
118k
•
673
zai-org/GLM-4.7
Text Generation
•
358B
•
Updated
Jan 29
•
102k
•
•
2.06k
MiniMaxAI/MiniMax-M2.1
Text Generation
•
229B
•
Updated
Feb 13
•
14.1k
•
•
1.36k
Load more