Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
6.8
TFLOPS
Pablo
p-ferrando
1
23
Follow
Gofar13's profile picture
1 follower
ยท
13 following
AI & ML interests
None yet
Recent Activity
reacted
to
pavle-scalably
's
post
with ๐ฅ
4 days ago
14 days serving https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4 to production agents on 2x RTX 5090 (vLLM 0.27, TP=2, 262K context, FP8 KV): 28,097 requests, 860.6M prompt tokens, 82.6% prefix-cache hit rate, TTFT p50 0.61 s, 0 engine errors. The observation: prefix cache, not throughput, decides whether a 27B model keeps up with agents. Mean request is 30,100 tokens in, 983 out, because every turn resends the whole session. Two flags mattered most: --max-num-seqs 12 (queue p95 went 9.4 s to 233.6 s past that) and --watermark 0.08 (preemptions 29 to 2). And thinking off for tool loops: 917 tokens in 11.7 s vs 11,170 in 144 s, same answer. Full config and counters: scalably.io/blog/qwen3-8-27b-nvfp4-rtx-5090-production Next we are preparing an 8x B300 node in an EU data center for open-weight serving. Which models or workloads are underserved for you?
liked
a model
7 days ago
prism-ml/Ternary-Bonsai-2-27B-gguf
liked
a model
15 days ago
deepseek-ai/DeepSeek-V4.1-Flash
View all activity
Organizations
None yet
spaces
1
Runtime error
Xatbot Canal Salut
๐ฅ
Answer health questions using Catalan health portal data
models
0
None public yet
datasets
0
None public yet