File size: 2,360 Bytes
80bda5e
 
e42a6ad
 
 
80bda5e
 
 
e42a6ad
80bda5e
e42a6ad
 
154f8cc
e42a6ad
80bda5e
 
d5efb77
80bda5e
e42a6ad
 
 
80bda5e
 
e42a6ad
154f8cc
e42a6ad
 
 
 
154f8cc
e42a6ad
154f8cc
e42a6ad
 
 
 
 
 
 
 
 
154f8cc
e42a6ad
 
 
 
154f8cc
 
 
e42a6ad
 
 
 
 
 
80bda5e
e42a6ad
80bda5e
e42a6ad
 
f9a0bcf
d5efb77
e42a6ad
 
 
154f8cc
e42a6ad
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---
library_name: premove-itn
language: [en]
base_model: [microsoft/deberta-v3-large]
tags: [inverse-text-normalization, speech-processing, voice-agents]
license: mit
---

# Premove ITN v0.2.0

This release updates the contextual candidate scorer used by `premove-itn`.
The deterministic Rust candidate generators and exact decoder are unchanged.

Load it through the matching Python package:

```python
from premove_itn import PremoveITN

itn = PremoveITN.from_pretrained(revision="v0.2.0")
print(itn.normalize("meet me at two thirty"))
# meet me at 02:30
```

## Training

The v0.1.0 scorer was adapted for three epochs on 7,440 controlled,
single-collision TIME-versus-identifier records. No replay or dual-collision
examples were used. Epoch 3 was selected using a separate 200-row development
set before any frozen test was inspected.

## Results

| Metric | v0.1.0 | v0.2.0 |
| --- | ---: | ---: |
| Context record | 51.2% | **86.2%** |
| Counterfactual pair | 6.4% | **72.4%** |
| Identifier | 29.6% | **98.0%** |
| Time | 72.8% | **74.4%** |
| Dual span | 45.0% | **96.0%** |
| Dual sentence exact | 16.0% | **92.0%** |
| Broad strict exact | 40.5% | **43.7%** |

The contextual and dual benchmarks are synthetic controlled evaluations. They
do not estimate production voice-agent accuracy. The broad benchmark gained 86
new exact rows and lost 38 previously exact rows. Localized losses were most
visible in ORDINAL, MONEY, URL, and DIGIT_SEQUENCE.

## Limitations

- TIME recall on the controlled single-collision test is 74.4%, substantially
  below the 98.0% identifier result.
- The model is English-only and requires the `premove-itn` candidate graph and
  decoder. It is not a generic Transformers model.
- The approximately 435.6M-parameter scorer has a large download and
  multi-second initialization cost.

## Release identity

- Artifact version: `v0.2.0`
- Required package version: `0.2.0`
- Hub repository: `premove-ai/premove-itn`
- Base model: `microsoft/deberta-v3-large`
- Base revision: `64a8c8eab3e352a784c658aef62be1662607476f`
- Source checkpoint SHA-256: `10a338d57b6d619f52259a58aca210ecb78f47de44baf0c25477a8f5b06b5d39`
- Model SHA-256: `0b6f36aa32311d0c495e1c5307d5e34463e52b4dc283ab9030bc2f54e1bf1152`

Source and evaluation evidence are available from
[`premove-ai/premove-itn`](https://github.com/premove-ai/premove-itn).