|
Download README.md from athrvk/gpu-sankhya: direct link, hf CLI and curl.
- Browser
- Download file 7.04 kB
-
https://huggingface.co/athrvk/gpu-sankhya/resolve/main/README.md
- Command line
-
hf download hf://athrvk/gpu-sankhya/README.md
-
curl -L -o README.md https://huggingface.co/athrvk/gpu-sankhya/resolve/main/README.md
7.04 kB
| license: mit | |
| language: | |
| - hi | |
| tags: | |
| - token-classification | |
| - onnx | |
| - char-cnn | |
| - hinglish | |
| - devanagari | |
| - indian-numbering | |
| - webgpu | |
| library_name: custom | |
| pipeline_tag: token-classification | |
| datasets: | |
| - athrvk/gpu-sankhya-gold | |
| # gpu-sankhya | |
| A small char-level CNN that extracts Indian informal number/currency | |
| shorthand -- Hinglish (romanised Hindi), Devanagari Hindi, Devanagari | |
| Marathi, Gujarati, and Indian-English amount phrases like `sava lakh`, `dedh crore`, `डेढ़ लाख`, | |
| `सवा करोड़`, `દોઢ લાખ`, `2.5L`, `20k`, `2-3 lakh` -- from free text, and turns each | |
| match into a clean numeric value via a deterministic arithmetic core (the | |
| model never predicts the value directly). | |
| - npm package (runtime, ships this model quantized inline): | |
| https://www.npmjs.com/package/gpu-sankhya | |
| - source / training code: | |
| https://github.com/athrvk/gpu-sankhya | |
| - live demo: https://huggingface.co/spaces/athrvk/gpu-sankhya-demo | |
| - gold evaluation sets: https://huggingface.co/datasets/athrvk/gpu-sankhya-gold | |
| This model card describes weights version `v0.8.0` (arch `v2`). | |
| ## Architecture | |
| A dilated/residual char-CNN over per-character embeddings: | |
| - embedding dim: 16 | |
| - conv channels: 48 | |
| - 5 conv layers: | |
| 1. kernel 5, dilation 1, residual no | |
| 2. kernel 3, dilation 1, residual yes | |
| 3. kernel 3, dilation 2, residual yes | |
| 4. kernel 3, dilation 4, residual yes | |
| 5. kernel 3, dilation 8, residual yes | |
| - vocab: 170 characters (union of every shipped language pack) | |
| - output classes: 120 (BIO span tag + semantic token class) | |
| - parameters: 40,475 | |
| A deterministic arithmetic core (not part of this model) then evaluates | |
| the decoded token sequence into a value: prefix semantics (sava = x1.25, | |
| dedh = x1.5, paune = subtract 1/4 from the next cardinal, ...), additive | |
| combination of descending units, multiplicative combination of ascending | |
| units, and range handling. | |
| ## Files | |
| | file | size | | |
| | --- | ---: | | |
| | `charset.json` | 1,425 bytes | | |
| | `classes.json` | 1,571 bytes | | |
| | `gold_metrics_json_float32.json` | 7,336 bytes | | |
| | `gold_metrics_json_int8.json` | 7,338 bytes | | |
| | `gold_metrics_torch.json` | 7,336 bytes | | |
| | `kaggle_metrics.json` | 26,135 bytes | | |
| | `matrix.md` | 872 bytes | | |
| | `sankhya.onnx` | 164,545 bytes | | |
| | `sankhya.pt` | 172,214 bytes | | |
| | `sankhya.weights.int8.json` | 60,720 bytes | | |
| | `sankhya.weights.json` | 405,469 bytes | | |
| - `sankhya.pt` -- torch checkpoint (vocab, classes, arch, channels, state_dict) | |
| - `sankhya.onnx` -- ONNX graph, for interop/inspection | |
| - `sankhya.weights.json` -- float32 weights, human-readable JSON | |
| - `sankhya.weights.int8.json` -- int8-quantized weights (what the npm | |
| package and this card's accuracy numbers use) | |
| - `charset.json`, `classes.json` -- standalone vocab/class tables | |
| - `gold_metrics_*.json` -- per-file + combined gold evaluation (torch, | |
| float32 JSON, int8 JSON) | |
| - `matrix.md`, `kaggle_metrics.json` -- training/sweep provenance from the | |
| Kaggle GPU run that produced this checkpoint | |
| ## Accuracy | |
| Evaluated with the int8-quantized weights (what ships in the npm package) | |
| against the hand-written gold sets (`gold_<lang>.jsonl` -- see the | |
| [dataset card](athrvk/gpu-sankhya-gold)): | |
| | gold set | examples | spans | value_acc | precision | recall | F1 | | |
| | --- | ---: | ---: | ---: | ---: | ---: | ---: | | |
| | gold.jsonl | 227 | 199 | 0.9598 | 0.9550 | 0.9598 | 0.9574 | | |
| | gold_deva.jsonl | 186 | 155 | 0.9677 | 0.9868 | 0.9677 | 0.9772 | | |
| | gold_mr.jsonl | 173 | 135 | 0.9852 | 0.9852 | 0.9852 | 0.9852 | | |
| | gold_gu.jsonl | 162 | 118 | 0.9322 | 0.9658 | 0.9576 | 0.9617 | | |
| | **combined** | 748 | 607 | **0.9621** | 0.9719 | 0.9671 | 0.9694 | | |
| Negatives (zero-gold-span examples): 170, false | |
| positives: 0 (0.00%). | |
| Per-category value accuracy (combined): | |
| | category | spans | value_acc | | |
| | --- | ---: | ---: | | |
| | digits | 145 | 0.9724 | | |
| | words | 462 | 0.9589 | | |
| | prefix | 104 | 0.9808 | | |
| | range | 49 | 0.9592 | | |
| | currency | 117 | 0.9658 | | |
| | multi_unit | 28 | 0.9643 | | |
| | symbol_unit | 50 | 1.0000 | | |
| | mixed_script | 3 | 1.0000 | | |
| | long | 16 | 0.8750 | | |
| ## Usage | |
| ### JavaScript (recommended -- ships this model quantized, no download) | |
| ```js | |
| import { parse } from "gpu-sankhya"; | |
| parse("sava lakh"); | |
| // [{ span: "sava lakh", value: 125000, unit: "lakh", ... }] | |
| ``` | |
| ### Python (numpy reference forward pass) | |
| ```python | |
| import numpy as np | |
| from sankhya import np_infer, decode, core, charset | |
| from sankhya.train import build_char_to_id | |
| weights_json = json.load(open("sankhya.weights.int8.json")) | |
| weights = np_infer.load_weights_int8_json(weights_json) | |
| char_to_id = build_char_to_id(weights_json["charset"]) | |
| unk = char_to_id.get("<unk>", 1) | |
| text = charset.normalize_text("sava lakh") | |
| ids = np_infer.pad_ids([char_to_id.get(c, unk) for c in text]) | |
| bio_logits, cls_logits = np_infer.forward(weights, np.array(ids)) | |
| n = len(text) | |
| bio_pred = bio_logits[:n].argmax(-1).tolist() | |
| cls_pred = cls_logits[:n].argmax(-1).tolist() | |
| bio_probs = np_infer.softmax(bio_logits[:n]).tolist() | |
| spans = decode.decode_spans(text, bio_pred, cls_pred, bio_probs=bio_probs) | |
| for span in spans: | |
| result = core.evaluate(span["tokens"]) | |
| print(text[span["start"]:span["end"]], result.value) | |
| ``` | |
| See `python/README.md` in the source repo for the full training/export | |
| pipeline and `sankhya.eval_gold` for a ready-made evaluation CLI. | |
| ## Training | |
| Trained 20 epochs on 200,000 synthetic examples generated from the | |
| `hi_latn` (romanised Hindi), `hi_deva` (Devanagari Hindi), `mr_deva` | |
| (Devanagari Marathi), and `gu_gujr` (Gujarati) grammar packs, mixed | |
| 0.32/0.26/0.21/0.21 with a 10% | |
| cross-pack share, plus out-of-vocab | |
| "unk noise" augmentation (emoji, CJK, Cyrillic, other symbols inserted as | |
| O-labelled context) so the `<unk>` embedding actually gets gradient | |
| signal. Batch size 128, lr 3e-3. Architecture and channel count (`v2`, | |
| 48 channels) were chosen by a config/seed sweep, picking the config with | |
| the highest mean int8 combined gold value_acc across >= 3 seeds (seed | |
| spread is +/-1-2 points), then the seed by val accuracy within a 0.005 tie | |
| band, then higher int8 combined gold F1, then lower negatives | |
| false-positive rate. Full recipe, sweep evidence, and reproduction | |
| commands: `python/README.md` in the source repo. | |
| ## Limitations | |
| - JavaScript string indices count UTF-16 code units, so astral characters | |
| (e.g. some emoji) occupy 2 code units -- span offsets from the JS | |
| runtime account for this, but consumers indexing raw strings themselves | |
| should be aware of it. | |
| - The model occasionally produces spurious spans on unfamiliar words near | |
| number-ish context (a measured trade-off from the out-of-vocab noise | |
| training -- see Accuracy above). | |
| - Long multi-term/mixed-numeral constructs and multi-number range phrases | |
| ("तीस पैंतीस हज़ार", "three n half lakh") are the weakest category | |
| (`long`/`range` value_acc above). | |
| - Scoped to Indian languages: currently Hinglish, Devanagari Hindi, | |
| Devanagari Marathi, and Gujarati only; other Indian languages are | |
| planned (see the source repo's roadmap). | |