hv-intent-router-word-2048

Parameters: 106 Γ— 2048 Γ— 2 = 434,176 trainable bits (434.2K)

Codebook size: 53 KB

Total model size: 71,863 bytes (~70 KB)

5-fold CV accuracy: ~73% (varies 71–76% across random seeds)

Random baseline: 20.0%

Training time: 0.6 seconds on a single CPU core

Dependencies: NumPy only


Model Card for hv-intent-router-word-2048

Model Description

A 70 KB hypervector intent classifier for 5-class routing (greeting, math, reminder, time, weather). The model uses word-level hyperdimensional computing with supervised word weighting and weighted k-nearest-neighbour voting. No neural networks, no gradient descent, no GPU.

The full model fits in L2 cache. Training completes in under one second. Inference is one similarity comparison per bank item.

Architecture

  • Tokens: Whitespace-split words. Vocabulary is 106 unique words drawn from the training phrases.
  • Encoding: Each word is assigned a 2048-bit bipolar hypervector. A phrase is encoded by summing weighted word hypervectors, then normalizing to unit length.
  • Supervised weights: Each word gets a weight in [0, 1] based on how discriminative it is across the 5 intents. Function words like "the" and "what" get weight near 0; discriminating words like "weather" and "calculate" get weight near 1. Words below 0.15 are dropped entirely.
  • Augmented bank: Each training phrase is expanded into 13 variants (1 original + 12 word-drop augments), giving a bank of 650 items.
  • Two codebooks: Two independent random codebooks, each encoding the full augmented bank.
  • Classifier: For each codebook, cosine similarity between query and every bank item. The top-7 nearest neighbours vote for their intent, weighted by similarity. Votes from both codebooks are summed. Argmax over 5 intents is the prediction.
  • Query augmentation: At inference, the query is dropped-augmented twice. All three variants vote.

Accuracy and Scaling

The model was tested at three codebook dimensions under identical conditions (5-fold CV, same augmentation, same weights, same classifier). Only D changed.

D Parameters Codebook size 5-fold CV Wall time
2,048 434,176 53 KB 71.0% 0.3s
20,480 4,341,760 530 KB 72.0% 1.2s
204,800 43,417,600 5.3 MB 72.0% 18.7s

The model's accuracy is flat across 100Γ— variation in parameter count. Going from 434 K to 43 M parameters buys one additional correct classification out of 100. Runtime scales linearly with D, as expected, but accuracy does not.

The same model run twice with different random seeds varies by 2–3 examples on the 50-phrase test set. The true accuracy of the D=2048 model is approximately 73% Β± 3%. The single-run number of 76% reported in an earlier version of this card was one sample from that distribution.

What this scaling result means

For this dataset, dimension is not the bottleneck. The model's errors come from phrases that share nearly all their vocabulary with a competing intent:

  • "when does the store close" (time) vs "what is the weather today" (weather) β€” both start with a wh-question word
  • "set an alarm for six" (reminder) vs "when is the meeting" (time) β€” both describe a scheduled event
  • "how hot is it" (weather) vs "how late is it" (time) β€” identical structure, different function word

Adding more dimensions cannot separate these phrases, because their word-level overlap is the same regardless of embedding size. Only additional training data or a different tokenisation scheme can help.

Where larger D does help

At D=2,048, the augmented bank of 650 items is near the measured capacity of a single hypervector (~64 items at 90% retrieval per hypervector, though here the bank uses one hypervector per bank item rather than bundling). Increasing D would be useful if the training set were 5–10Γ— larger, because the number of distinct discriminative patterns would exceed what D=2,048 can represent.

The recommendation: use D=2,048 for datasets up to ~100 training phrases. Above that, increase D proportionally to the number of distinct phrases per intent.

Training Data

Intent Phrases
greeting 10
math 10
reminder 10
time 10
weather 10
Total 50

Training phrases are short English commands, e.g. "what is the weather today", "remind me to call mom", "what is two plus two".

Evaluation

Model Accuracy Size Training Time
Random codebook 20.0% 53 KB 0s
This model (5-fold CV) ~73% 70 KB 0.6s
This model at 100Γ— D 72.0% 5.3 MB 18.7s
DistilBERT (reference, 66M params) ~95% 250 MB hours

This model reaches ~73% with roughly 4000Γ— fewer parameters than a distilled transformer. It fits comfortably in an always-on memory region of any modern CPU.

History of this model

Version Method 5-fold CV
v1 Character-level, single memory, GA-evolved codebook 44.0%
v2 Word-level, single memory 52.0%
v3 Word + word-drop augmentation 68.0%
v4 Word + augmentation + supervised word weights 73–76%
v4 at D=20,480 Same, 10Γ— parameters 72.0%
v4 at D=204,800 Same, 100Γ— parameters 72.0%

Each row is measured under identical 5-fold CV. The jump from 44% to 73% came from changing the representation (character to word, then adding weights and augmentation). The jump from 73% to 72% is noise.

Intended Use

  • Edge intent routing β€” classify short commands on microcontrollers, DSPs, or always-on co-processors where a transformer cannot fit.
  • Pre-filtering β€” route requests to the correct downstream model before invoking it, saving compute on out-of-scope queries.
  • Few-shot classification β€” train in under a second on a new label set without GPU or autograd.
  • On-device personalisation β€” retrain per user with a handful of example phrases.

Limitations

  • Closed vocabulary. The 106-word codebook covers only words that appear in the training data. Out-of-vocabulary words are silently dropped.
  • Small training set. 50 phrases across 5 intents. The 5-fold CV estimate has a 95% confidence interval of roughly 62–84%. The true generalisation accuracy is likely 71–76%.
  • No semantics. The model learns word-level co-occurrence, not meaning. Paraphrases that share no words with the training set will not be classified correctly.
  • Bag of words. Word order is discarded. "what time is it" and "is it time" produce similar encodings.
  • No handling of negation. "Is it not sunny" and "is it sunny" produce nearly identical encodings.
  • Parameter scaling does not improve accuracy on this dataset. See the scaling table above.

How to Use

from hv_intent import load_v2, predict

model = load_v2("zeechimp/zee")
predict(model, "what is the weather today")   # -> "weather"
predict(model, "remind me to call mom")       # -> "reminder"
predict(model, "calculate ten times three")   # -> "math"
Downloads last month
53
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results

  • 5-fold CV Accuracy (mean over seeds) on 5-class intent routing
    self-reported
    73.000
  • Random Baseline on 5-class intent routing
    self-reported
    20.000