gpu-sankhya / README.md
athrvk's picture
Upload README.md with huggingface_hub
4e306b1 verified
|
Raw History Blame Contribute Delete
7.04 kB
---
license: mit
language:
- hi
tags:
- token-classification
- onnx
- char-cnn
- hinglish
- devanagari
- indian-numbering
- webgpu
library_name: custom
pipeline_tag: token-classification
datasets:
- athrvk/gpu-sankhya-gold
---
# gpu-sankhya
A small char-level CNN that extracts Indian informal number/currency
shorthand -- Hinglish (romanised Hindi), Devanagari Hindi, Devanagari
Marathi, Gujarati, and Indian-English amount phrases like `sava lakh`, `dedh crore`, `डेढ़ लाख`,
`सवा करोड़`, `દોઢ લાખ`, `2.5L`, `20k`, `2-3 lakh` -- from free text, and turns each
match into a clean numeric value via a deterministic arithmetic core (the
model never predicts the value directly).
- npm package (runtime, ships this model quantized inline):
https://www.npmjs.com/package/gpu-sankhya
- source / training code:
https://github.com/athrvk/gpu-sankhya
- live demo: https://huggingface.co/spaces/athrvk/gpu-sankhya-demo
- gold evaluation sets: https://huggingface.co/datasets/athrvk/gpu-sankhya-gold
This model card describes weights version `v0.8.0` (arch `v2`).
## Architecture
A dilated/residual char-CNN over per-character embeddings:
- embedding dim: 16
- conv channels: 48
- 5 conv layers:
1. kernel 5, dilation 1, residual no
2. kernel 3, dilation 1, residual yes
3. kernel 3, dilation 2, residual yes
4. kernel 3, dilation 4, residual yes
5. kernel 3, dilation 8, residual yes
- vocab: 170 characters (union of every shipped language pack)
- output classes: 120 (BIO span tag + semantic token class)
- parameters: 40,475
A deterministic arithmetic core (not part of this model) then evaluates
the decoded token sequence into a value: prefix semantics (sava = x1.25,
dedh = x1.5, paune = subtract 1/4 from the next cardinal, ...), additive
combination of descending units, multiplicative combination of ascending
units, and range handling.
## Files
| file | size |
| --- | ---: |
| `charset.json` | 1,425 bytes |
| `classes.json` | 1,571 bytes |
| `gold_metrics_json_float32.json` | 7,336 bytes |
| `gold_metrics_json_int8.json` | 7,338 bytes |
| `gold_metrics_torch.json` | 7,336 bytes |
| `kaggle_metrics.json` | 26,135 bytes |
| `matrix.md` | 872 bytes |
| `sankhya.onnx` | 164,545 bytes |
| `sankhya.pt` | 172,214 bytes |
| `sankhya.weights.int8.json` | 60,720 bytes |
| `sankhya.weights.json` | 405,469 bytes |
- `sankhya.pt` -- torch checkpoint (vocab, classes, arch, channels, state_dict)
- `sankhya.onnx` -- ONNX graph, for interop/inspection
- `sankhya.weights.json` -- float32 weights, human-readable JSON
- `sankhya.weights.int8.json` -- int8-quantized weights (what the npm
package and this card's accuracy numbers use)
- `charset.json`, `classes.json` -- standalone vocab/class tables
- `gold_metrics_*.json` -- per-file + combined gold evaluation (torch,
float32 JSON, int8 JSON)
- `matrix.md`, `kaggle_metrics.json` -- training/sweep provenance from the
Kaggle GPU run that produced this checkpoint
## Accuracy
Evaluated with the int8-quantized weights (what ships in the npm package)
against the hand-written gold sets (`gold_<lang>.jsonl` -- see the
[dataset card](athrvk/gpu-sankhya-gold)):
| gold set | examples | spans | value_acc | precision | recall | F1 |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| gold.jsonl | 227 | 199 | 0.9598 | 0.9550 | 0.9598 | 0.9574 |
| gold_deva.jsonl | 186 | 155 | 0.9677 | 0.9868 | 0.9677 | 0.9772 |
| gold_mr.jsonl | 173 | 135 | 0.9852 | 0.9852 | 0.9852 | 0.9852 |
| gold_gu.jsonl | 162 | 118 | 0.9322 | 0.9658 | 0.9576 | 0.9617 |
| **combined** | 748 | 607 | **0.9621** | 0.9719 | 0.9671 | 0.9694 |
Negatives (zero-gold-span examples): 170, false
positives: 0 (0.00%).
Per-category value accuracy (combined):
| category | spans | value_acc |
| --- | ---: | ---: |
| digits | 145 | 0.9724 |
| words | 462 | 0.9589 |
| prefix | 104 | 0.9808 |
| range | 49 | 0.9592 |
| currency | 117 | 0.9658 |
| multi_unit | 28 | 0.9643 |
| symbol_unit | 50 | 1.0000 |
| mixed_script | 3 | 1.0000 |
| long | 16 | 0.8750 |
## Usage
### JavaScript (recommended -- ships this model quantized, no download)
```js
import { parse } from "gpu-sankhya";
parse("sava lakh");
// [{ span: "sava lakh", value: 125000, unit: "lakh", ... }]
```
### Python (numpy reference forward pass)
```python
import numpy as np
from sankhya import np_infer, decode, core, charset
from sankhya.train import build_char_to_id
weights_json = json.load(open("sankhya.weights.int8.json"))
weights = np_infer.load_weights_int8_json(weights_json)
char_to_id = build_char_to_id(weights_json["charset"])
unk = char_to_id.get("<unk>", 1)
text = charset.normalize_text("sava lakh")
ids = np_infer.pad_ids([char_to_id.get(c, unk) for c in text])
bio_logits, cls_logits = np_infer.forward(weights, np.array(ids))
n = len(text)
bio_pred = bio_logits[:n].argmax(-1).tolist()
cls_pred = cls_logits[:n].argmax(-1).tolist()
bio_probs = np_infer.softmax(bio_logits[:n]).tolist()
spans = decode.decode_spans(text, bio_pred, cls_pred, bio_probs=bio_probs)
for span in spans:
result = core.evaluate(span["tokens"])
print(text[span["start"]:span["end"]], result.value)
```
See `python/README.md` in the source repo for the full training/export
pipeline and `sankhya.eval_gold` for a ready-made evaluation CLI.
## Training
Trained 20 epochs on 200,000 synthetic examples generated from the
`hi_latn` (romanised Hindi), `hi_deva` (Devanagari Hindi), `mr_deva`
(Devanagari Marathi), and `gu_gujr` (Gujarati) grammar packs, mixed
0.32/0.26/0.21/0.21 with a 10%
cross-pack share, plus out-of-vocab
"unk noise" augmentation (emoji, CJK, Cyrillic, other symbols inserted as
O-labelled context) so the `<unk>` embedding actually gets gradient
signal. Batch size 128, lr 3e-3. Architecture and channel count (`v2`,
48 channels) were chosen by a config/seed sweep, picking the config with
the highest mean int8 combined gold value_acc across >= 3 seeds (seed
spread is +/-1-2 points), then the seed by val accuracy within a 0.005 tie
band, then higher int8 combined gold F1, then lower negatives
false-positive rate. Full recipe, sweep evidence, and reproduction
commands: `python/README.md` in the source repo.
## Limitations
- JavaScript string indices count UTF-16 code units, so astral characters
(e.g. some emoji) occupy 2 code units -- span offsets from the JS
runtime account for this, but consumers indexing raw strings themselves
should be aware of it.
- The model occasionally produces spurious spans on unfamiliar words near
number-ish context (a measured trade-off from the out-of-vocab noise
training -- see Accuracy above).
- Long multi-term/mixed-numeral constructs and multi-number range phrases
("तीस पैंतीस हज़ार", "three n half lakh") are the weakest category
(`long`/`range` value_acc above).
- Scoped to Indian languages: currently Hinglish, Devanagari Hindi,
Devanagari Marathi, and Gujarati only; other Indian languages are
planned (see the source repo's roadmap).