File size: 2,360 Bytes
80bda5e e42a6ad 80bda5e e42a6ad 80bda5e e42a6ad 154f8cc e42a6ad 80bda5e d5efb77 80bda5e e42a6ad 80bda5e e42a6ad 154f8cc e42a6ad 154f8cc e42a6ad 154f8cc e42a6ad 154f8cc e42a6ad 154f8cc e42a6ad 80bda5e e42a6ad 80bda5e e42a6ad f9a0bcf d5efb77 e42a6ad 154f8cc e42a6ad | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 | ---
library_name: premove-itn
language: [en]
base_model: [microsoft/deberta-v3-large]
tags: [inverse-text-normalization, speech-processing, voice-agents]
license: mit
---
# Premove ITN v0.2.0
This release updates the contextual candidate scorer used by `premove-itn`.
The deterministic Rust candidate generators and exact decoder are unchanged.
Load it through the matching Python package:
```python
from premove_itn import PremoveITN
itn = PremoveITN.from_pretrained(revision="v0.2.0")
print(itn.normalize("meet me at two thirty"))
# meet me at 02:30
```
## Training
The v0.1.0 scorer was adapted for three epochs on 7,440 controlled,
single-collision TIME-versus-identifier records. No replay or dual-collision
examples were used. Epoch 3 was selected using a separate 200-row development
set before any frozen test was inspected.
## Results
| Metric | v0.1.0 | v0.2.0 |
| --- | ---: | ---: |
| Context record | 51.2% | **86.2%** |
| Counterfactual pair | 6.4% | **72.4%** |
| Identifier | 29.6% | **98.0%** |
| Time | 72.8% | **74.4%** |
| Dual span | 45.0% | **96.0%** |
| Dual sentence exact | 16.0% | **92.0%** |
| Broad strict exact | 40.5% | **43.7%** |
The contextual and dual benchmarks are synthetic controlled evaluations. They
do not estimate production voice-agent accuracy. The broad benchmark gained 86
new exact rows and lost 38 previously exact rows. Localized losses were most
visible in ORDINAL, MONEY, URL, and DIGIT_SEQUENCE.
## Limitations
- TIME recall on the controlled single-collision test is 74.4%, substantially
below the 98.0% identifier result.
- The model is English-only and requires the `premove-itn` candidate graph and
decoder. It is not a generic Transformers model.
- The approximately 435.6M-parameter scorer has a large download and
multi-second initialization cost.
## Release identity
- Artifact version: `v0.2.0`
- Required package version: `0.2.0`
- Hub repository: `premove-ai/premove-itn`
- Base model: `microsoft/deberta-v3-large`
- Base revision: `64a8c8eab3e352a784c658aef62be1662607476f`
- Source checkpoint SHA-256: `10a338d57b6d619f52259a58aca210ecb78f47de44baf0c25477a8f5b06b5d39`
- Model SHA-256: `0b6f36aa32311d0c495e1c5307d5e34463e52b4dc283ab9030bc2f54e1bf1152`
Source and evaluation evidence are available from
[`premove-ai/premove-itn`](https://github.com/premove-ai/premove-itn).
|