Router Multidomain

Kicaulah AI Downloads license base params format context Demo

The dispatcher for Kicaulah AI - reads a message, picks one of five specialists, and hands over its system prompt. 135M parameters, 0.4s per message on CPU.

Try the live demo โ†’ ยท browse all six system prompts and copy them straight into any model.


What this is

A five-model agent system needs something to decide who answers. That something is this: a distilbert classifier over five labels.

                  user message
                        |
             [ Router Multidomain ]  <-- you are here
                        |
   therapist   health   education   cybersec   coding
      |          |          |           |          |
   model-    model-     model-      model-    model-
   therapist  health   education   cybersec   coding

It is deliberately tiny and CPU-only. Routing is not an interesting problem and does not deserve a GPU.


Quick start

from transformers import pipeline

pipe = pipeline("text-classification", model="Kicaulah/router-multidomain")

for text in [
    "I'm so tired of life lately",
    "my chest feels tight and I can't breathe properly",
    "Why does the sky look red at sunset",
    "I keep getting texts asking for my bank password",
    "write me a bash script to rename files in bulk",
]:
    r = pipe(text)[0]
    print(f"{text[:44]:46} -> {r['label']:11} {r['score']:.3f}")
I'm so tired of life lately                    -> therapist   0.891
my chest feels tight and I can't breathe ...   -> health      0.908
Why does the sky look red at sunset            -> health      0.842
I keep getting texts asking for my bank pass... -> therapist   0.803
write me a bash script to rename files in bulk -> coding      0.375

That last pair is the honest picture: it is right most of the time, and the two errors are real. See the breakdown below rather than the headline.


Accuracy

Measurement Value
Held-out split (30 messages) 0.8333
5-fold cross-validation (all 150) 0.733 ยฑ 0.070
Fold range 0.60 โ€“ 0.80
F1 (weighted, held-out) 0.8357
Random baseline (5 classes) 0.200
Inference cost ~0.4s on CPU

Why the CV number is the real number

A 30-message test split cannot distinguish 0.80 from 0.87. That is a two-sample difference, which is noise. scripts/eval_router_cv.py runs stratified 5-fold CV over all 150 examples, so the headline figure has a confidence interval behind it:

$ python scripts/eval_router_cv.py 5 10
fold 1/5: accuracy 0.7667  f1 0.7610
fold 2/5: accuracy 0.8000  f1 0.7914
fold 3/5: accuracy 0.7667  f1 0.7660
fold 4/5: accuracy 0.6000  f1 0.5872
fold 5/5: accuracy 0.7333  f1 0.7181
accuracy  mean 0.7330  std 0.0700
95% CI on accuracy: +/- 0.061

On 16 realistic out-of-distribution queries it scored 69%, which is consistent.

Where it still fails

Straight from router_eval_results.json:

Input Predicted Should be
I keep getting weird texts asking for my bank password therapist cybersec
two-factor authentication keeps failing on my phone therapist cybersec
Why does the sky look red at sunset health education
can you explain compound interest therapist education

Two structural problems, both fixable with more data rather than a bigger model:

  1. cybersec is the hardest label. Security questions phrased as personal trouble ("I keep getting texts asking for my password") read as emotional, so they land on therapist.
  2. health and therapist share vocabulary. Physical and emotional complaints both use words like worry, can't cope, stressed.

If you are putting this in production

Do not trust a bare argmax. Gate on confidence and fall back to the user:

result = pipe(message)[0]
if result["score"] < 0.60:
    reply = "I want to make sure I get this right - is this about your health, "
    reply += "your code, or something you want to talk through?"
else:
    domain = result["label"]

Design note: why the base model is multilingual

The original spec called for distilbert-base-uncased. That model is English-only, and its tokenizer shreds other languages into arbitrary subword fragments, so those patterns are effectively invisible to the model. On a dataset this size the result is not reliable.

distilbert-base-multilingual-cased is 135M params, covers 50+ languages, and costs the same to train. Real users type in whatever language they think in, and this routes them correctly instead of shrugging.

To follow the original spec, set BASE_MODEL in scripts/01_train_router.py.


Dataset

  • 150 English examples, 5 domains x 30
  • Everyday phrasing, written by hand, not scraped
  • 80/20 stratified split, seed 42
  • Regenerate: python scripts/make_router_dataset.py

Seven early examples were rewritten because they were genuinely ambiguous across domains ("What is artificial intelligence?" belongs to education, coding, and cybersec). Training on items with no correct answer teaches noise.


Limitations

  • Classification only. It does not answer anything. It picks a destination.
  • Mixed-topic messages get misrouted. "I feel stressed and my React app is broken" has no single right answer here.
  • Small dataset. 150 examples demonstrates that routing works. It is not a production classifier.
  • English-tuned. The base model is multilingual, but the fine-tune data is English, so non-English routing is untested.

Ethical note

The router itself gives no advice of any kind. It hands messages to specialists that each carry their own disclaimer.

One design point worth stating: the crisis guardrail in Kicaulah/model-therapist does not depend on this router. It runs on the raw text before the router is consulted, because this router misroutes "kms" and "suicidal" to education. A safety check that depends on an ~73%-accurate classifier is not a safety check.



Live demo

huggingface.co/spaces/Kicaulah/Kicaulah-AI-Demo

Browse all six system prompts with a worked example for each, and copy them straight into any instruct model. No download required.

The Kicaulah AI ecosystem

Repo Role What it does
Kicaulah/router-multidomain Router classifies the message, picks a specialist โ† you are here
Kicaulah/model-therapist Therapist warm, empathetic, never judges
Kicaulah/model-health Health calm, informative, names the red flags
Kicaulah/model-education Education patient teacher, everyday analogies
Kicaulah/model-cybersec CyberSec senior engineer, defensive only
Kicaulah/model-coding Coding pragmatic senior dev, blunt

License

Apache-2.0. Base model distilbert-base-multilingual-cased is Apache-2.0.

Downloads last month
66
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Kicaulah/router-multidomain

Finetuned
(446)
this model

Space using Kicaulah/router-multidomain 1