--- license: mit language: - hi tags: - token-classification - onnx - char-cnn - hinglish - devanagari - indian-numbering - webgpu library_name: custom pipeline_tag: token-classification datasets: - athrvk/gpu-sankhya-gold --- # gpu-sankhya A small char-level CNN that extracts Indian informal number/currency shorthand -- Hinglish (romanised Hindi), Devanagari Hindi, Devanagari Marathi, Gujarati, and Indian-English amount phrases like `sava lakh`, `dedh crore`, `डेढ़ लाख`, `सवा करोड़`, `દોઢ લાખ`, `2.5L`, `20k`, `2-3 lakh` -- from free text, and turns each match into a clean numeric value via a deterministic arithmetic core (the model never predicts the value directly). - npm package (runtime, ships this model quantized inline): https://www.npmjs.com/package/gpu-sankhya - source / training code: https://github.com/athrvk/gpu-sankhya - live demo: https://huggingface.co/spaces/athrvk/gpu-sankhya-demo - gold evaluation sets: https://huggingface.co/datasets/athrvk/gpu-sankhya-gold This model card describes weights version `v0.8.0` (arch `v2`). ## Architecture A dilated/residual char-CNN over per-character embeddings: - embedding dim: 16 - conv channels: 48 - 5 conv layers: 1. kernel 5, dilation 1, residual no 2. kernel 3, dilation 1, residual yes 3. kernel 3, dilation 2, residual yes 4. kernel 3, dilation 4, residual yes 5. kernel 3, dilation 8, residual yes - vocab: 170 characters (union of every shipped language pack) - output classes: 120 (BIO span tag + semantic token class) - parameters: 40,475 A deterministic arithmetic core (not part of this model) then evaluates the decoded token sequence into a value: prefix semantics (sava = x1.25, dedh = x1.5, paune = subtract 1/4 from the next cardinal, ...), additive combination of descending units, multiplicative combination of ascending units, and range handling. ## Files | file | size | | --- | ---: | | `charset.json` | 1,425 bytes | | `classes.json` | 1,571 bytes | | `gold_metrics_json_float32.json` | 7,336 bytes | | `gold_metrics_json_int8.json` | 7,338 bytes | | `gold_metrics_torch.json` | 7,336 bytes | | `kaggle_metrics.json` | 26,135 bytes | | `matrix.md` | 872 bytes | | `sankhya.onnx` | 164,545 bytes | | `sankhya.pt` | 172,214 bytes | | `sankhya.weights.int8.json` | 60,720 bytes | | `sankhya.weights.json` | 405,469 bytes | - `sankhya.pt` -- torch checkpoint (vocab, classes, arch, channels, state_dict) - `sankhya.onnx` -- ONNX graph, for interop/inspection - `sankhya.weights.json` -- float32 weights, human-readable JSON - `sankhya.weights.int8.json` -- int8-quantized weights (what the npm package and this card's accuracy numbers use) - `charset.json`, `classes.json` -- standalone vocab/class tables - `gold_metrics_*.json` -- per-file + combined gold evaluation (torch, float32 JSON, int8 JSON) - `matrix.md`, `kaggle_metrics.json` -- training/sweep provenance from the Kaggle GPU run that produced this checkpoint ## Accuracy Evaluated with the int8-quantized weights (what ships in the npm package) against the hand-written gold sets (`gold_.jsonl` -- see the [dataset card](athrvk/gpu-sankhya-gold)): | gold set | examples | spans | value_acc | precision | recall | F1 | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | gold.jsonl | 227 | 199 | 0.9598 | 0.9550 | 0.9598 | 0.9574 | | gold_deva.jsonl | 186 | 155 | 0.9677 | 0.9868 | 0.9677 | 0.9772 | | gold_mr.jsonl | 173 | 135 | 0.9852 | 0.9852 | 0.9852 | 0.9852 | | gold_gu.jsonl | 162 | 118 | 0.9322 | 0.9658 | 0.9576 | 0.9617 | | **combined** | 748 | 607 | **0.9621** | 0.9719 | 0.9671 | 0.9694 | Negatives (zero-gold-span examples): 170, false positives: 0 (0.00%). Per-category value accuracy (combined): | category | spans | value_acc | | --- | ---: | ---: | | digits | 145 | 0.9724 | | words | 462 | 0.9589 | | prefix | 104 | 0.9808 | | range | 49 | 0.9592 | | currency | 117 | 0.9658 | | multi_unit | 28 | 0.9643 | | symbol_unit | 50 | 1.0000 | | mixed_script | 3 | 1.0000 | | long | 16 | 0.8750 | ## Usage ### JavaScript (recommended -- ships this model quantized, no download) ```js import { parse } from "gpu-sankhya"; parse("sava lakh"); // [{ span: "sava lakh", value: 125000, unit: "lakh", ... }] ``` ### Python (numpy reference forward pass) ```python import numpy as np from sankhya import np_infer, decode, core, charset from sankhya.train import build_char_to_id weights_json = json.load(open("sankhya.weights.int8.json")) weights = np_infer.load_weights_int8_json(weights_json) char_to_id = build_char_to_id(weights_json["charset"]) unk = char_to_id.get("", 1) text = charset.normalize_text("sava lakh") ids = np_infer.pad_ids([char_to_id.get(c, unk) for c in text]) bio_logits, cls_logits = np_infer.forward(weights, np.array(ids)) n = len(text) bio_pred = bio_logits[:n].argmax(-1).tolist() cls_pred = cls_logits[:n].argmax(-1).tolist() bio_probs = np_infer.softmax(bio_logits[:n]).tolist() spans = decode.decode_spans(text, bio_pred, cls_pred, bio_probs=bio_probs) for span in spans: result = core.evaluate(span["tokens"]) print(text[span["start"]:span["end"]], result.value) ``` See `python/README.md` in the source repo for the full training/export pipeline and `sankhya.eval_gold` for a ready-made evaluation CLI. ## Training Trained 20 epochs on 200,000 synthetic examples generated from the `hi_latn` (romanised Hindi), `hi_deva` (Devanagari Hindi), `mr_deva` (Devanagari Marathi), and `gu_gujr` (Gujarati) grammar packs, mixed 0.32/0.26/0.21/0.21 with a 10% cross-pack share, plus out-of-vocab "unk noise" augmentation (emoji, CJK, Cyrillic, other symbols inserted as O-labelled context) so the `` embedding actually gets gradient signal. Batch size 128, lr 3e-3. Architecture and channel count (`v2`, 48 channels) were chosen by a config/seed sweep, picking the config with the highest mean int8 combined gold value_acc across >= 3 seeds (seed spread is +/-1-2 points), then the seed by val accuracy within a 0.005 tie band, then higher int8 combined gold F1, then lower negatives false-positive rate. Full recipe, sweep evidence, and reproduction commands: `python/README.md` in the source repo. ## Limitations - JavaScript string indices count UTF-16 code units, so astral characters (e.g. some emoji) occupy 2 code units -- span offsets from the JS runtime account for this, but consumers indexing raw strings themselves should be aware of it. - The model occasionally produces spurious spans on unfamiliar words near number-ish context (a measured trade-off from the out-of-vocab noise training -- see Accuracy above). - Long multi-term/mixed-numeral constructs and multi-number range phrases ("तीस पैंतीस हज़ार", "three n half lakh") are the weakest category (`long`/`range` value_acc above). - Scoped to Indian languages: currently Hinglish, Devanagari Hindi, Devanagari Marathi, and Gujarati only; other Indian languages are planned (see the source repo's roadmap).