YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

ServingStudio token corpora

Recorded per-token MoE routes for routing: corpus in ServingStudio Sim. Each directory is one capture: manifest.json (dimensions + FNV-1a checksum), routes.u16 (little-endian u16, C order [token, layer, top_k]), and provenance.json. Reference a file by commit, never by branch:

token_corpus_file: hf://UW-SyFI/servingstudio-corpora@<commit-sha>/glm52_nvfp4_mtp5/manifest.json

glm52_nvfp4_mtp5

  • Model: nvidia/GLM-5.2-NVFP4 @ aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa
  • Deployment: vLLM, B200, TP4 EP4 DP1, MTP num_speculative_tokens=5, fp8 KV
  • Workload: alignment pack glm52_nvfp4_b200_spec5, case 07_quadrant_interference_c48 (192 requests, enwik9 prompts, concurrency 48)
  • Captured by the profile_corpus (profile_kind: token_corpus) pass, Slurm job 952, 2026-09-23
  • 307,008 accepted generated tokens x 76 layers (75 body MoE layers + the MTP layer) x top-8 of 256 experts
  • Scope: accepted generated tokens only; no prompt or rejected-draft routes

glm53_nvfp4_dflash2

  • Model: incoai/GLM-5.3-NVFP4 @ 54e52520606f96b3d9fc84088ad22882a61648ac, proposer incoai/GLM-5.3-DFlash2 @ 425aa615ce320caac34400208b30808c8f14f76c
  • Deployment: vLLM (fork 79838a5a, v0.28 base), B200, TP4 EP4 DP1, DFlash2 num_speculative_tokens=7, fp8 KV, max_num_seqs 64
  • Workload: trace/diverse_100.csv (100 requests, enwik9 prompts, concurrency 64)
  • Captured by a one-off capture_corpus.py (non-streaming /v1/completions routed_experts), Slurm job 823, 2026-09-22
  • 135,144 accepted generated tokens x 75 layers (the body MoE layers; the target runs no MTP layer) x top-8 of 256 experts
  • Scope: accepted generated tokens only; no prompt or rejected-draft routes

glm53_flash_fp8_tp4_ep4

  • Model: zai-org/GLM-5.3-Flash @ 03eb5366286afd40d2221b1d9c63a6dd1ba4832e (FP8 block checkpoint)
  • Deployment: vLLM (fork 3f667d7e, v0.28 base), B200, TP4 EP4 DP1, no speculative decoding, fp8 KV, max-model-len 8192
  • Workload: 256 independent requests (input 256/1024/2048/4096 x output 128/256, 32 each), enwik9 prompts, saturated at concurrency 32
  • Captured by the profile_kind: token_corpus pass, Slurm job 1071, 2026-09-24
  • 48,896 accepted generated tokens x 42 layers (the MoE layers 3-44) x top-8 of 288 experts
  • Scope: accepted generated tokens only; no prompt routes
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support