Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
6.8
TFLOPS
Pablo
p-ferrando
1
23
Follow
wbf8d797ec1df44c67's profile picture
Gofar13's profile picture
2 followers
·
14 following
AI & ML interests
None yet
Recent Activity
reacted
to
pavle-scalably
's
post
with 🔥
12 days ago
14 days serving https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4 to production agents on 2x RTX 5090 (vLLM 0.27, TP=2, 262K context, FP8 KV): 28,097 requests, 860.6M prompt tokens, 82.6% prefix-cache hit rate, TTFT p50 0.61 s, 0 engine errors. The observation: prefix cache, not throughput, decides whether a 27B model keeps up with agents. Mean request is 30,100 tokens in, 983 out, because every turn resends the whole session. Two flags mattered most: --max-num-seqs 12 (queue p95 went 9.4 s to 233.6 s past that) and --watermark 0.08 (preemptions 29 to 2). And thinking off for tool loops: 917 tokens in 11.7 s vs 11,170 in 144 s, same answer. Full config and counters: scalably.io/blog/qwen3-8-27b-nvfp4-rtx-5090-production Next we are preparing an 8x B300 node in an EU data center for open-weight serving. Which models or workloads are underserved for you?
liked
a model
15 days ago
prism-ml/Ternary-Bonsai-2-27B-gguf
liked
a model
23 days ago
deepseek-ai/DeepSeek-V4.1-Flash
View all activity
Organizations
None yet
p-ferrando
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
15 days ago
prism-ml/Ternary-Bonsai-2-27B-gguf
Text Generation
•
27B
•
Updated
8 days ago
•
3.97M
•
2.36k
liked
a model
23 days ago
deepseek-ai/DeepSeek-V4.1-Flash
Image-Text-to-Text
•
763B
•
Updated
2 days ago
•
788k
•
•
4.02k
liked
2 models
about 2 months ago
Qwen/Qwen3.8-27B
Image-Text-to-Text
•
28B
•
Updated
Aug 14
•
6.9M
•
•
16.8k
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation
•
304B
•
Updated
Aug 1
•
4.48M
•
•
4.01k
liked
a model
2 months ago
Nanbeige/Nanbeige4.2-3B
Text Generation
•
4B
•
Updated
22 days ago
•
78.5k
•
809
liked
2 models
4 months ago
sapientinc/HRM-Text-1B
Text Generation
•
1B
•
Updated
29 days ago
•
19.7k
•
824
LiquidAI/LFM2.5-8B-A1B
Text Generation
•
8B
•
Updated
Aug 24
•
32.4k
•
779
liked
2 models
5 months ago
FINAL-Bench/Darwin-28B-Opus
Text Generation
•
28B
•
Updated
Jul 23
•
92
•
•
43
openbmb/MiniCPM-V-4.6
Image-Text-to-Text
•
1B
•
Updated
Aug 17
•
287k
•
1.23k
liked
2 models
6 months ago
Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text
•
36B
•
Updated
Apr 24
•
3.31M
•
•
2.91k
FINAL-Bench/Darwin-31B-Opus
Text Generation
•
31B
•
Updated
Aug 11
•
103
•
•
65
liked
a model
7 months ago
mistralai/Mistral-Small-4-119B-2603
119B
•
Updated
Jul 15
•
54.3k
•
430
liked
8 models
9 months ago
baichuan-inc/Baichuan-M3-235B
Text Generation
•
235B
•
Updated
Feb 9
•
468
•
104
mistralai/Mistral-Large-3-675B-Instruct-2512
Updated
Jul 15
•
1.69k
•
252
mistralai/Ministral-3-14B-Instruct-2512
14B
•
Updated
Jul 15
•
251k
•
330
deepseek-ai/DeepSeek-V3.2-Speciale
Text Generation
•
685B
•
Updated
Dec 1, 2025
•
3.56k
•
726
tencent/Youtu-LLM-2B
Text Generation
•
2B
•
Updated
Feb 24
•
11.6k
•
231
LiquidAI/LFM2.5-1.2B-Instruct
Text Generation
•
1B
•
Updated
Aug 24
•
119k
•
673
zai-org/GLM-4.7
Text Generation
•
358B
•
Updated
Jan 29
•
102k
•
•
2.06k
MiniMaxAI/MiniMax-M2.1
Text Generation
•
229B
•
Updated
Feb 13
•
13.8k
•
•
1.36k
Load more