MiniLM Example Router

A browser-ready example-based router, using frozen MiniLM embeddings and normalized mean vectors per route. Build custom routes from your own examples in the interactive editor. No GPU, API key, backend or gradient training is needed in the editor.

This is an integration and evaluated prototype-head release, not a newly trained sentence encoder. All encoder weights come unchanged from sentence-transformers/all-MiniLM-L6-v2, revision 1110a243fdf4706b3f48f1d95db1a4f5529b4d41. We export normalized embeddings as ONNX and fit route vectors from labeled examples. This is not a standard Transformers pipeline checkpoint. Use the browser worker and router modules in web/, or implement cosine ranking against the provided centroids.

What is included

  • encoder.onnx:45,305,375 bytes, 384-dimensional unit embeddings, maximum256 tokens; large matrices stored FP16 and cast to FP32 for computation. Runtime memory is not necessarily halved. Inputs:int64 input_ids,attention_mask, token_type_ids; output:embeddings. Attention-mask mean pooling + L2 normalization.
  • router_seed11.json:150-intent CLINC router with8 support examples per intent. Seed11 was fixed first, not chosen for best test accuracy. Seed22/33 are also supplied for reproducibility. These are fixed-label benchmark examples; custom routes in the editor start from the user's own labeled examples.
  • Native tokenizer assets, browser integration source, tests and evaluation.

Use it

Open the editor, replace the fictional sample routes with your labels and examples (one per line), then build and try a message. Export JSON to keep the router; importing restores route vectors for inference. Exports omit original example text, so editing an imported router requires re-entering examples. Derived vectors and route names may still be sensitive: keep exported files private when appropriate. No text is sent to an inference service. Model/runtime downloads require connectivity; browser caching is not an offline guarantee.

For a static-host integration see web/README.md. Router JSON is bound to embedding model identifier minilm-l6-v2-fp16-v1; mixing encoders or pooling rules makes centroids invalid. Core ranking is dot product between unit message embedding and unit route centroid. Scores are similarities, not calibrated probabilities. No default review threshold is applied to custom routes.

Deploy exported routes in Python

Use the CPU Python guide to load your browser-exported JSON in a script or service, without PyTorch or Transformers. Explicit local files support offline inference. Verification covers a real browser export, numeric fixtures, list inputs, invalid files and thresholds.

Evaluation

Frozen encoder; no contrastive fine-tuning. CLINC150 plus, revision 155b9c710419136e17307b80d0a13e68cd46b4ec,8 train examples per in-scope intent, seeds11/22/33. Deduplicated training, removed label conflicts and normalized train-text overlaps from validation/test. Never select the best seed.

Seed Test accuracy, in-scope In-scope rejection OOS rejection
11 87.68% 4.42% 77.2%
22 88.48% 4.71% 75.9%
33 87.11% 4.25% 75.9%

Mean closed-set accuracy87.76%, on4,498 in-scope examples after2 train-overlap removals. OOS test has1,000 examples. Rejection uses each seed's5th percentile of in-scope validation max-cosine. Rejection figures above are original FP32 encoder results, separate from closed-set accuracy; they do not establish performance for a different route set. Thresholds in benchmark JSON are research settings, not recommended defaults for custom routes. ONNX rejection parity has not been separately measured.

Validation prototype mean87.75%; same-support fixed-C1 logistic heads averaged 87.32%; label-name cosine70.04%. These have different supervision budgets: prototypes/logistic receive8 labeled examples per intent, label cosine gets names only. No superiority over SetFit fine-tuning is claimed; SetFit was not run. Pretraining overlap with CLINC is unknown, so this is not proof of unseen-data or novel-intent generalization. Benchmark intents are predefined.

The compressed ONNX encoder plus centroids refitted from ONNX support vectors preserves in-scope accuracy across all seeds; top-label agreement with FP32 is 99.9818%,99.9818%,100% over evaluated queries (including forced labels for OOS). Maximum coordinate error0.000220. Full reports in evaluation/ distinguish native CPU ONNX checks from browser smoke tests.

Limitations

English short-message routing only. Long inputs truncated to256 tokens. Similar routes, poor examples, multilingual inputs and domain shift can degrade results. Nearest-route classification always chooses a route unless an optional threshold requests review. An OOS threshold cannot guarantee unknown-intent detection. The sample demo routes are fictional and are not part of the benchmark. Do not use routing labels as authorization for actions or claim security protection. Evaluate with held-out examples and retain a human fallback where needed.

Provenance and licenses

Encoder/tokenizer:Apache2.0, upstream attribution retained. Code and fitted prototype-head files:MIT. CLINC data:CC-BY3.0, Stefan Larson et al.(2019), An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction. Dataset. No raw CLINC corpus redistributed here; reports include source hashes/support hashes. Respect upstream dataset terms when reproducing. Research scripts use cluster-specific source paths that must be adjusted to your artifact directories.

Release tests and artifact checksums are included. No download counters from our own verification should be interpreted as organic adoption.

Deployment configuration

Both current loaders consume root config.json for the embedding contract and pinned artifact identity. Python checks encoder/tokenizer hashes; the browser checks the encoder hash. Older pinned clients remain usable. Configuration loads follow the Hub default download-count convention; counts include internal checks and are not unique-user or organic-adoption measurements.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Future-Labs/minilm-example-router

Dataset used to train Future-Labs/minilm-example-router

Space using Future-Labs/minilm-example-router 1