MoS DFlash Routers

This repository contains three experimental request-level router snapshots used by the MoS/DFlash studies. A router classifies one prompt and selects one of five draft experts:

0 code
1 math
2 factual_qa
3 creative_writing
4 general

These artifacts are not token-level routers, standalone language models, or complete serving systems. They require hidden states from the exact target model and feature contract recorded beside each weight file.

Release layout

request-level/
  offline-v2/qwen3-8b/
  intraining-sidecar/qwen3-8b/
  c3/qwen3-4b/
registry/
schema/
verification/
manifests/
Release Feature contract Head
offline-v2/qwen3-8b layers 1/9/17/25/33; mean, max, and last pooling; stored z-score statistics LayerNorm(61440) β†’ Linear(512) β†’ GELU β†’ Linear(5)
intraining-sidecar/qwen3-8b layers 1/9/17/25/33; mean pooling LayerNorm(20480) β†’ Linear(512) β†’ GELU β†’ Linear(5)
c3/qwen3-4b layers 1/9/17/25/33; mean pooling LayerNorm(12800) β†’ Linear(512) β†’ GELU β†’ Linear(5)

The offline-v2 source was a PyTorch checkpoint. This release converts only its tensor state into safetensors and does not publish the pickle container.

Loading

Install the two runtime dependencies:

pip install -r requirements.txt

Then load a release with the included helper:

from pathlib import Path
from router_loader import load_offline_v2, load_sidecar

root = Path("request-level/offline-v2/qwen3-8b")
head, feature_mean, feature_std, config = load_offline_v2(
    root / "router_head.safetensors",
    root / "router_config.json",
)
head.eval()

load_sidecar handles either in-training sidecar release. Feature extraction is intentionally not hidden inside the loader: callers must implement the recorded tokenizer, prompt template, target-layer, and pooling contract exactly.

The offline-v2 deployment contract additionally zeros z-scored dimensions whose stored standard deviation is at most 2e-6, then clamps the remaining values to [-10, 10]. Use prepare_offline_features from router_loader.py.

Verification boundary

The release manifest binds every public file to SHA-256 and byte size. The included validation report covers:

  • exact tensor equality for the offline-v2 weights-only conversion;
  • safetensors header, key, shape, dtype, and finite-forward checks;
  • removal of machine-local paths and private per-sample material.

It does not claim task-level accuracy reproduction on a new inference backend. Exact tokenizer, target revision, prompt formatting, hidden-state indexing, and numeric dtype still need end-to-end parity testing.

Raw prompts, hidden-state feature tensors, per-sample routing records, training logs, and raw pickle checkpoints are not included.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support