MoS-DFlash-Routers / README.md
ryan-0608's picture
hf3-request-level-v1: metadata, configs, code, and registry (#1)
29d1f68
|
Raw
History Blame Contribute Delete
2.98 kB
---
license: mit
inference: false
tags:
- routing
- request-level-routing
- speculative-decoding
- dflash
- mixture-of-speculators
- qwen3
---
# MoS DFlash Routers
This repository contains three experimental **request-level** router snapshots
used by the MoS/DFlash studies. A router classifies one prompt and selects one
of five draft experts:
```text
0 code
1 math
2 factual_qa
3 creative_writing
4 general
```
These artifacts are not token-level routers, standalone language models, or
complete serving systems. They require hidden states from the exact target
model and feature contract recorded beside each weight file.
## Release layout
```text
request-level/
offline-v2/qwen3-8b/
intraining-sidecar/qwen3-8b/
c3/qwen3-4b/
registry/
schema/
verification/
manifests/
```
| Release | Feature contract | Head |
|---|---|---|
| `offline-v2/qwen3-8b` | layers 1/9/17/25/33; mean, max, and last pooling; stored z-score statistics | LayerNorm(61440) β†’ Linear(512) β†’ GELU β†’ Linear(5) |
| `intraining-sidecar/qwen3-8b` | layers 1/9/17/25/33; mean pooling | LayerNorm(20480) β†’ Linear(512) β†’ GELU β†’ Linear(5) |
| `c3/qwen3-4b` | layers 1/9/17/25/33; mean pooling | LayerNorm(12800) β†’ Linear(512) β†’ GELU β†’ Linear(5) |
The offline-v2 source was a PyTorch checkpoint. This release converts only its
tensor state into safetensors and does not publish the pickle container.
## Loading
Install the two runtime dependencies:
```bash
pip install -r requirements.txt
```
Then load a release with the included helper:
```python
from pathlib import Path
from router_loader import load_offline_v2, load_sidecar
root = Path("request-level/offline-v2/qwen3-8b")
head, feature_mean, feature_std, config = load_offline_v2(
root / "router_head.safetensors",
root / "router_config.json",
)
head.eval()
```
`load_sidecar` handles either in-training sidecar release. Feature extraction
is intentionally not hidden inside the loader: callers must implement the
recorded tokenizer, prompt template, target-layer, and pooling contract
exactly.
The offline-v2 deployment contract additionally zeros z-scored dimensions
whose stored standard deviation is at most `2e-6`, then clamps the remaining
values to `[-10, 10]`. Use `prepare_offline_features` from
`router_loader.py`.
## Verification boundary
The release manifest binds every public file to SHA-256 and byte size. The
included validation report covers:
- exact tensor equality for the offline-v2 weights-only conversion;
- safetensors header, key, shape, dtype, and finite-forward checks;
- removal of machine-local paths and private per-sample material.
It does **not** claim task-level accuracy reproduction on a new inference
backend. Exact tokenizer, target revision, prompt formatting, hidden-state
indexing, and numeric dtype still need end-to-end parity testing.
Raw prompts, hidden-state feature tensors, per-sample routing records, training
logs, and raw pickle checkpoints are not included.