percent-nlu

Live demo · GitHub · npm · In production: icupu.com Percentage Calculator

A 1.16M-parameter natural-language understanding model that reads a percentage question in English, Portuguese or Spanish and returns the right calculator tool call with every number in the right role:

"how much is an 18% tip on $64?"       → percent_of(rate=18, base=64)            = 11.52
"quanto é 30% de desconto em 250?"      → percent_of(rate=30, base=250)           = 75
"el alquiler subió de 1.450 a 1.595"    → percent_increase(initial=1450, final=1595) = 10%
"investi 100 e agora tenho 60"          → percent_decrease(initial=100, final=60) = 40%
"200 minus 15%"                         → apply_decrease(base=200, rate=15)       = 170
"what percentage of the earth is water" → out_of_scope

It is built to run inside the browser on any phone: 1.2 MB of int8 weights, a dependency-free JavaScript runtime, ~1–5 ms per question on CPU. The model never does arithmetic — it only picks the tool and the roles; the numbers are parsed deterministically and the result is computed by plain code.

Tools (output space)

tool meaning arguments
percent_of X% of Y rate, base
what_percent X is what % of Y part, whole
percent_increase % increase from A to B initial, final
percent_decrease % decrease from A to B initial, final
apply_increase Y plus X% base, rate
apply_decrease Y minus X% base, rate
— not a percentage calculation out_of_scope

Each answer comes with a status: call (confident, safe to execute), low_confidence (best guess, ask the user to confirm), incomplete (intent understood, a number is missing — e.g. "20% of my salary") or out_of_scope.

How it works

  1. Normalizer (rules, per-language lexicon) — finds numbers written as digits (1,234.56, 1.234,56, 5k, $, €, R$), as words ("one hundred and twenty-five", "doscientos cincuenta y tres mil", "dois mil e trezentos") and percent markers (%, percent, por cento, por ciento), then masks them: what is [N0] % of [N1].
  2. Transformer encoder (2 layers, d=128, pre-LN) over word embeddings + hashed character 3–5-grams (FNV-1a) + small per-number features (has %, has currency, magnitude, largest/smallest…). Two heads: intent, and a role for each [Nk].
  3. Constrained decoding — each tool's required roles are assigned to distinct numbers; extra numbers get none.
  4. Calibrated confidence + gate — temperature-scaled probabilities; status = call only when confidence ≥ τ (τ = 0.774, chosen on the dev set). Below τ the call is still returned as low_confidence.

Language is given by the caller (recommended: the page locale) or detected from function words.

Usage

Browser / Node (JavaScript, no dependencies)

import { NLU } from "./web/percent-nlu.js";           // ES module
const nlu = await NLU.load("./web");                   // fetches manifest.json + weights.bin (1.2 MB)
const r = nlu.predict("how much is an 18% tip on $64?", "en");   // lang: "en" | "pt" | "es" | undefined (auto)
// r.status → "call", r.tool → "percent_of", r.args → {rate: 18, base: 64}, r.result → 11.52, r.confidence, r.timing

Try it live in the demo Space.

Or install it: npm install percent-nlu (npm).

Python

# pip install "percent-nlu @ git+https://github.com/EmanoelV/percent-nlu#subdirectory=python"   (model bundled)
# or, from a clone of this Hugging Face repo, run Python in its root folder
import percent_nlu
p = percent_nlu.load()
p("cuánto es el 15% de 80?", lang="es")
# {'lang': 'es', 'status': 'call', 'tool': 'percent_of', 'args': {'rate': 15.0, 'base': 80.0}, 'result': 12.0, ...}

Evaluation

Exact = intent and every role correct. OOS accepted = out-of-scope questions wrongly treated as calculations. Gate precision / coverage = accuracy of the answers with status = call, and the share of answerable questions that get it.

set n exact OOS accepted gate precision gate coverage
gold v0 (pt-BR) 242 97.6% 6.0% 98.7% 95.7%
gold v1 (pt-BR, hard) 112 90.3% 0.0% 96.1% 81.7%
gold v1 (English) 118 95.7% 10.0% 100.0% 91.3%
gold v1 (Español) 118 94.6% 0.0% 98.8% 92.4%
synthetic test, pt-BR (unseen templates) 3812 90.5% 7.0% 98.1% 82.6%
synthetic test, English (unseen templates) 3815 95.1% 7.4% 99.3% 85.1%
synthetic test, Español (unseen templates) 3812 98.2% 19.4% 96.4% 91.2%
  • Gold sets are questions written one by one (with and without %, inverted order, context before the numbers, slang), with LLM assistance, before the training templates of each language, and never used for training or model selection.
  • Synthetic test uses templates (sentence patterns) that never appear in training: the split is by template, not by example.
  • Size: 1,160,462 parameters (vocabulary 2,522 words + 4,096 hashed n-gram buckets) · CPU latency p50 0.6 ms (PyTorch, batch 1, Apple M1) · JavaScript: ~1.4 ms on an M1, ~4.6 ms on a mid-range Android phone.

Training

  • Data: ~115k synthetic questions (≈38k per language) generated from 1,288 Portuguese, 754 English, 723 Spanish sentence templates written with LLM assistance, with varied number formats, currencies, number words, typos, missing accents, and missing % signs; out-of-scope questions with numbers (battery %, polls, unit conversions, plain arithmetic, device commands). No user data is used for training.
  • Recipe: 10 epochs, AdamW (lr 3e-3, cosine), batch 512, label smoothing 0.05, word dropout 0.1 (random words → <unk>, so the model does not rely on a single cue word), EMA of weights (0.99), temperature calibration on dev.
  • Run: percent-nlu 1.0.0.

Limitations

  • Only the six tools above; at most 4 numbers per question; no chained calculations ("20% off and then 10% tax").
  • The gold sets come from the same authoring process as the templates (LLM-assisted); real users will phrase things differently. Use the status field: execute call, confirm low_confidence.
  • Number parsing is rule-based. Not supported yet: fractions (1/4), "half a percent", decimals in words ("seven point five"), ranges. In Spanish, 1,500 is read as one thousand five hundred (groups of exactly 3 digits are thousands for both . and ,); 1,5 and 1.5 are 1.5.
  • Language detection is a simple function-word count; pass lang when you know it.
  • percent_increase vs percent_decrease for sentences without a direction word ("I had 100, now 60") is decided by the numbers; the tools also return a direction_mismatch warning when the numbers contradict the wording.

Files

path content
config.json architecture, vocabulary, calibration (τ, temperatures), intents/roles/tools, number lexicons
model.safetensors fp32 weights (PyTorch)
web/manifest.json, web/weights.bin browser bundle: int8 weights (per-row scales) + lexicons
web/percent-nlu.js JavaScript runtime (normalizer, tokenizer, Transformer, decoding, tools)
percent_nlu/ Python inference code (normalizer, featurizer, model, decoding, tools)
eval/metrics.json evaluation metrics
LICENSE MIT
Downloads last month
24
Safetensors
Model size
1.16M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using EmanoelV/percent-nlu 1