Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
ShoaibSSM
/
a100-qwen-api-runtime
Like
0
License:
mit
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
a100-qwen-api-runtime
/
scripts
13.6 kB
Ctrl+K
Ctrl+K
1 contributor
History:
4 commits
ShoaibSSM
Add optional GLM-4.7-Flash Q8 and GPT-OSS-120B MXFP4 aliases
48f42a4
verified
8 days ago
bench-throughput.sh
Safe
659 Bytes
Add pinned A100 Qwen API runtime and benchmark scripts
8 days ago
bench.py
Safe
3.07 kB
Add lazy two-model router and detached ngrok setup
8 days ago
check_threads.sh
Safe
443 Bytes
Bound llama.cpp HTTP and compute threads; enable four slots
8 days ago
install.sh
Safe
3.3 kB
Add pinned A100 Qwen API runtime and benchmark scripts
8 days ago
router.sh
Safe
3.01 kB
Add optional GLM-4.7-Flash Q8 and GPT-OSS-120B MXFP4 aliases
8 days ago
serve.sh
Safe
2.3 kB
Bound llama.cpp HTTP and compute threads; enable four slots
8 days ago
start.sh
Safe
561 Bytes
Add lazy two-model router and detached ngrok setup
8 days ago
stop.sh
Safe
309 Bytes
Add pinned A100 Qwen API runtime and benchmark scripts
8 days ago