Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
jabbatheduck
/
ninfer-ext-models
Like
2
Image-Text-to-Text
Trellis
NInfer
English
Chinese
code
ninfer-ext
exl3
nvfp4
Mixture of Experts
expert-offload
blackwell
quantization
cuda
text-generation
vision
speculative-decoding
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Deploy
Copy to bucket
new
main
ninfer-ext-models
159 GB
Ctrl+K
Ctrl+K
1 contributor
History:
25 commits
jabbatheduck
Qwen3.8-Flash-Next: label prefill with ngram residency; note mapped is ~1.7x faster
737e67e
verified
2 days ago
qwen3.8-flash-next
Qwen3.8-Flash-Next: label prefill with ngram residency; note mapped is ~1.7x faster
2 days ago
.gitattributes
Safe
2.13 kB
Add Qwen3.8-Flash-Next NVFP4 (ModelOpt, 119 GB, 4 shards)
3 days ago
README.md
10.1 kB
Qwen3.8-Flash-Next: label prefill with ngram residency; note mapped is ~1.7x faster
2 days ago
qwen3_8_27b_exl3_3p5bpw.ninfer
15.3 GB
xet
Rebuild with the indexed draft head and the fork chat template
5 days ago
qwen3_8_27b_exl3_3p5bpw.ninfer.conversion.json
491 kB
Rebuild with the indexed draft head and the fork chat template
5 days ago
qwen3_8_27b_exl3_4bpw.ninfer
16.8 GB
xet
Rebuild with the indexed draft head and the fork chat template
5 days ago
qwen3_8_27b_exl3_4bpw.ninfer.conversion.json
491 kB
Rebuild with the indexed draft head and the fork chat template
5 days ago