percent-nlu
Live demo · GitHub · npm · In production: icupu.com Percentage Calculator
A 1.16M-parameter natural-language understanding model that reads a percentage question in English, Portuguese or Spanish and returns the right calculator tool call with every number in the right role:
"how much is an 18% tip on $64?" → percent_of(rate=18, base=64) = 11.52
"quanto é 30% de desconto em 250?" → percent_of(rate=30, base=250) = 75
"el alquiler subió de 1.450 a 1.595" → percent_increase(initial=1450, final=1595) = 10%
"investi 100 e agora tenho 60" → percent_decrease(initial=100, final=60) = 40%
"200 minus 15%" → apply_decrease(base=200, rate=15) = 170
"what percentage of the earth is water" → out_of_scope
It is built to run inside the browser on any phone: 1.2 MB of int8 weights, a dependency-free JavaScript runtime, ~1–5 ms per question on CPU. The model never does arithmetic — it only picks the tool and the roles; the numbers are parsed deterministically and the result is computed by plain code.
Tools (output space)
| tool | meaning | arguments |
|---|---|---|
percent_of |
X% of Y | rate, base |
what_percent |
X is what % of Y | part, whole |
percent_increase |
% increase from A to B | initial, final |
percent_decrease |
% decrease from A to B | initial, final |
apply_increase |
Y plus X% | base, rate |
apply_decrease |
Y minus X% | base, rate |
| — | not a percentage calculation | out_of_scope |
Each answer comes with a status: call (confident, safe to execute), low_confidence (best guess, ask the user to
confirm), incomplete (intent understood, a number is missing — e.g. "20% of my salary") or out_of_scope.
How it works
- Normalizer (rules, per-language lexicon) — finds numbers written as digits (
1,234.56,1.234,56,5k,$,€,R$), as words ("one hundred and twenty-five", "doscientos cincuenta y tres mil", "dois mil e trezentos") and percent markers (%, percent, por cento, por ciento), then masks them:what is [N0] % of [N1]. - Transformer encoder (2 layers, d=128, pre-LN) over word embeddings + hashed character 3–5-grams (FNV-1a) + small
per-number features (has %, has currency, magnitude, largest/smallest…). Two heads: intent, and a role for each
[Nk]. - Constrained decoding — each tool's required roles are assigned to distinct numbers; extra numbers get
none. - Calibrated confidence + gate — temperature-scaled probabilities;
status = callonly when confidence ≥ τ (τ = 0.774, chosen on the dev set). Below τ the call is still returned aslow_confidence.
Language is given by the caller (recommended: the page locale) or detected from function words.
Usage
Browser / Node (JavaScript, no dependencies)
import { NLU } from "./web/percent-nlu.js"; // ES module
const nlu = await NLU.load("./web"); // fetches manifest.json + weights.bin (1.2 MB)
const r = nlu.predict("how much is an 18% tip on $64?", "en"); // lang: "en" | "pt" | "es" | undefined (auto)
// r.status → "call", r.tool → "percent_of", r.args → {rate: 18, base: 64}, r.result → 11.52, r.confidence, r.timing
Try it live in the demo Space.
Or install it: npm install percent-nlu (npm).
Python
# pip install "percent-nlu @ git+https://github.com/EmanoelV/percent-nlu#subdirectory=python" (model bundled)
# or, from a clone of this Hugging Face repo, run Python in its root folder
import percent_nlu
p = percent_nlu.load()
p("cuánto es el 15% de 80?", lang="es")
# {'lang': 'es', 'status': 'call', 'tool': 'percent_of', 'args': {'rate': 15.0, 'base': 80.0}, 'result': 12.0, ...}
Evaluation
Exact = intent and every role correct. OOS accepted = out-of-scope questions wrongly treated as calculations.
Gate precision / coverage = accuracy of the answers with status = call, and the share of answerable questions that get it.
| set | n | exact | OOS accepted | gate precision | gate coverage |
|---|---|---|---|---|---|
| gold v0 (pt-BR) | 242 | 97.6% | 6.0% | 98.7% | 95.7% |
| gold v1 (pt-BR, hard) | 112 | 90.3% | 0.0% | 96.1% | 81.7% |
| gold v1 (English) | 118 | 95.7% | 10.0% | 100.0% | 91.3% |
| gold v1 (Español) | 118 | 94.6% | 0.0% | 98.8% | 92.4% |
| synthetic test, pt-BR (unseen templates) | 3812 | 90.5% | 7.0% | 98.1% | 82.6% |
| synthetic test, English (unseen templates) | 3815 | 95.1% | 7.4% | 99.3% | 85.1% |
| synthetic test, Español (unseen templates) | 3812 | 98.2% | 19.4% | 96.4% | 91.2% |
- Gold sets are questions written one by one (with and without
%, inverted order, context before the numbers, slang), with LLM assistance, before the training templates of each language, and never used for training or model selection. - Synthetic test uses templates (sentence patterns) that never appear in training: the split is by template, not by example.
- Size: 1,160,462 parameters (vocabulary 2,522 words + 4,096 hashed n-gram buckets) · CPU latency p50 0.6 ms (PyTorch, batch 1, Apple M1) · JavaScript: ~1.4 ms on an M1, ~4.6 ms on a mid-range Android phone.
Training
- Data: ~115k synthetic questions (≈38k per language) generated from 1,288 Portuguese, 754 English, 723 Spanish sentence templates written with
LLM assistance, with varied number formats, currencies, number words, typos, missing accents, and missing
%signs; out-of-scope questions with numbers (battery %, polls, unit conversions, plain arithmetic, device commands). No user data is used for training. - Recipe: 10 epochs, AdamW (lr 3e-3, cosine), batch 512, label smoothing 0.05, word dropout 0.1 (random words
→
<unk>, so the model does not rely on a single cue word), EMA of weights (0.99), temperature calibration on dev. - Run:
percent-nlu 1.0.0.
Limitations
- Only the six tools above; at most 4 numbers per question; no chained calculations ("20% off and then 10% tax").
- The gold sets come from the same authoring process as the templates (LLM-assisted); real users will phrase things
differently. Use the
statusfield: executecall, confirmlow_confidence. - Number parsing is rule-based. Not supported yet: fractions (1/4), "half a percent", decimals in words
("seven point five"), ranges. In Spanish,
1,500is read as one thousand five hundred (groups of exactly 3 digits are thousands for both.and,);1,5and1.5are 1.5. - Language detection is a simple function-word count; pass
langwhen you know it. percent_increasevspercent_decreasefor sentences without a direction word ("I had 100, now 60") is decided by the numbers; the tools also return adirection_mismatchwarning when the numbers contradict the wording.
Files
| path | content |
|---|---|
config.json |
architecture, vocabulary, calibration (τ, temperatures), intents/roles/tools, number lexicons |
model.safetensors |
fp32 weights (PyTorch) |
web/manifest.json, web/weights.bin |
browser bundle: int8 weights (per-row scales) + lexicons |
web/percent-nlu.js |
JavaScript runtime (normalizer, tokenizer, Transformer, decoding, tools) |
percent_nlu/ |
Python inference code (normalizer, featurizer, model, decoding, tools) |
eval/metrics.json |
evaluation metrics |
LICENSE |
MIT |
- Downloads last month
- 24