|
Download README.md from NagaYu/downstep: direct link, hf CLI and curl.
- Browser
- Download file 9.4 kB
-
https://huggingface.co/NagaYu/downstep/resolve/main/README.md
- Command line
-
hf download hf://NagaYu/downstep/README.md
-
curl -L -o README.md https://huggingface.co/NagaYu/downstep/resolve/main/README.md
9.4 kB
| language: | |
| - ja | |
| license: apache-2.0 | |
| library_name: downstep | |
| pipeline_tag: token-classification | |
| tags: | |
| - text-to-speech | |
| - japanese | |
| - pitch-accent | |
| - prosody | |
| - accent-phrase | |
| - openjtalk | |
| - voicevox | |
| - style-bert-vits2 | |
| - calibration | |
| - selective-prediction | |
| - onnx | |
| metrics: | |
| - exact_match | |
| - f1 | |
| # Downstep — Japanese pitch accent for compounds your dictionary has never seen | |
| **Getting 東京大学 right is a lookup. The hard part is 新型感染症対策本部.** | |
| Downstep is a 35M-parameter encoder that predicts the **accent nucleus** and **accent-phrase | |
| boundaries** of a Japanese noun compound from its morphemes. It reports a calibrated | |
| confidence and, below a threshold, **abstains** — returning the existing front end's answer | |
| unchanged — so it can be added to an OpenJTalk-based TTS pipeline without replacing it. | |
| Code, benchmark builder and evaluation: **https://github.com/NagaYu/downstep** | |
| --- | |
| ## What this model is for, and what it is not for | |
| **Use it for** the fallback path: compounds that are *not* entries in your TTS dictionary, | |
| where a rule-based front end has to compose the accent from parts. | |
| **Do not use it to replace dictionary lookup.** On compounds that *are* in the dictionary this | |
| model is clearly worse than OpenJTalk (81.3% vs 96.4%), and worse than the dictionary itself | |
| (100% by definition). That is not a defect — the dictionary is right and there is nothing to | |
| add. The shipped adapter tries the dictionary first for exactly this reason. | |
| This is a **text-only** model. It reads a surface string and a kana reading; it has never seen | |
| speech. Accent models trained on thousands of hours of audio are solving a different problem. | |
| --- | |
| ## Results | |
| Evaluated on 5,414 held-out compounds, scored identically in both conditions. Gold accents | |
| come from NAIST-jdic and UniDic 2.1.2. Metric: accent-nucleus exact match. | |
| | system | compound **in** dictionary | compound **not** in dictionary | | |
| |---|---|---| | |
| | OpenJTalk (rule-based) | **96.4%** | 40.6% | | |
| | dictionary lookup only | 100.0% | 40.6% (falls back) | | |
| | **Downstep** | 81.3% | **81.3%** | | |
| | **Downstep + abstention** | **96.3%** | **81.3%** | | |
| **+40.7 points** on dictionary-unlisted compounds. On the *hard* subset — compounds whose | |
| final morpheme never appeared in training — **73.3% vs 40.9%**, so the gain is not memorisation. | |
| OpenJTalk also wrongly splits an unlisted compound into two accent phrases 23.3% of the time; | |
| Downstep does so 2.2% of the time. | |
| The "not in dictionary" condition is produced by **ablating the compound's own entry** at query | |
| time so the gold stays real. It is a held-out-entry simulation, not an observation of genuinely | |
| novel compounds — see Limitations. | |
| **Calibration.** Temperature scaling fitted on dev: ECE **0.048 → 0.011** (T = 1.498). | |
| **n-best.** top-1 81.3%, top-3 **98.0%**, top-5 99.7%. On words whose dictionaries attest | |
| *more than one* accent, top-1 is 61.2% but top-3 is 98.0%. | |
| **Speed.** 2.0 ms/word (PyTorch, CPU). ONNX int8: 1.3 ms at batch 1, ~3× faster than fp32. | |
| --- | |
| ## Files | |
| | file | what it is | size | | |
| |---|---|---| | |
| | `model.pt` | PyTorch checkpoint (config + weights) | 134 MB | | |
| | `vocab.json` | **required** — the learned feature id maps | 12 KB | | |
| | `downstep.onnx` | ONNX fp32, opset 17, dynamic batch and time | 134 MB | | |
| | `downstep.int8.onnx` | ONNX int8 dynamic quantisation | 34 MB | | |
| | `calibrator.json`, `calibration.json` | temperature scaler + dev-chosen thresholds | small | | |
| `vocab.json` is not optional. Feature id spaces are learned from the training corpus, so a | |
| mismatched vocab does not raise — it silently predicts nonsense. | |
| **The int8 build's decisions are identical** to fp32 (nucleus argmax and boundary sign agree | |
| 100%), but its logits deviate enough to fail a 5e-02 tolerance. Since the abstention gate | |
| thresholds a *confidence*, re-calibrate before using int8 with a threshold rather than | |
| inheriting the fp32 one. | |
| There is **no GGUF**. llama.cpp dispatches on `general.architecture` and has no `downstep` | |
| entry, and its token-ids-in / vocabulary-logits-out interface does not describe this model, | |
| which takes 13 parallel feature streams and returns three per-mora logit vectors. ONNX is the | |
| export that actually runs. | |
| --- | |
| ## Usage | |
| ```bash | |
| pip install "downstep[all] @ git+https://github.com/NagaYu/downstep" | |
| ``` | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| from downstep.adapters.openjtalk import DropInAdapter | |
| path = snapshot_download("NagaYu/downstep") | |
| adapter = DropInAdapter.load(path) # dictionary first, model for the rest | |
| # VOICEVOX-shaped accent phrases | |
| adapter.accent_phrases("新型感染症対策本部を設置した") | |
| # NJD features — same keys as pyopenjtalk.run_frontend(), swap it in directly | |
| adapter.njd_features("新型感染症対策本部を設置した") | |
| # HTS full-context labels | |
| adapter.full_context_labels("新型感染症対策本部を設置した") | |
| ``` | |
| For **VOICEVOX**, substitute `adapter.accent_phrases(text)` where the engine builds its | |
| `AudioQuery` accent phrases. For **Style-Bert-VITS2**, point its `pyopenjtalk.run_frontend` | |
| call at `adapter.njd_features`. | |
| **A trap worth knowing.** In HTS full-context labels the accent field `F2` writes 平板 (flat, | |
| no fall) as `F2 == F1`, *not* as 0 — so `F2 == F1` is ambiguous between 平板 and 尾高 (nucleus | |
| on the final mora). Writing a literal 0 there silently corrupts every flat-accent word, and | |
| reading `F2` naively scores OpenJTalk at 70% where the truth is 94%. The adapter converts in | |
| both directions and raises `AmbiguousF2Error` rather than guessing. | |
| --- | |
| ## Training | |
| - **Data:** 71,544 compounds from NAIST-jdic, each decomposing into ≥2 morphemes whose | |
| readings compose to the compound's reading under 連濁 (rendaku) and 促音化. Decomposition | |
| uses fugashi + UniDic — deliberately a *different* dictionary from the gold, so the parts | |
| handed to the model are not derived from the answer. | |
| - **Architecture:** 10-layer pre-LN transformer encoder, d_model 512, 8 heads, d_ff 2048. | |
| 13 per-mora integer feature streams (mora identity, position, morpheme boundary, word | |
| origin, each component's own dictionary accent and sandhi class, …) summed as embeddings. | |
| Three per-mora heads: boundary, nucleus, heiban. | |
| - **Objective:** boundary BCE + a per-phrase softmax over (each mora position, plus a | |
| no-nucleus option). That single softmax per phrase is the training-time counterpart of the | |
| decoder's hard constraint that an accent phrase carries **at most one nucleus**. | |
| - **Run:** 12 epochs, batch 256, AdamW 3e-4 with cosine decay, 82 minutes on Apple MPS. | |
| - 25% of training items are **constructed** two-phrase examples (two compounds joined by a | |
| case particle), because every real benchmark compound is a single accent phrase and the | |
| boundary head would otherwise see only negative examples. | |
| --- | |
| ## Limitations | |
| - **No human annotation exists anywhere in this project.** Every label is a dictionary entry. | |
| The compounds this model is built for are precisely the ones no dictionary contains, and | |
| their accent can only be settled by asking native speakers. The repository ships the full | |
| annotation protocol, multi-annotator format, agreement statistics and an extractor that | |
| emits novel compounds **unlabeled** — but nothing here has been checked by a person. | |
| - **The unlisted condition is simulated.** Held-out-entry ablation on lexicalised compounds | |
| is what makes real gold available; it is not an observation of genuinely novel compounds. | |
| - **Per-word non-degradation is expensive.** At full coverage the model makes 256 of 5,414 | |
| words worse while making 2,462 better (~10:1). Driving individual regressions to zero | |
| requires abstaining on ~93% of items, which scores 42.2% instead of 81.3%. In aggregate the | |
| combined system is at or above the baseline at every coverage level; individual words are | |
| not. Choose the threshold deliberately. | |
| - **Proper nouns are 66% of the benchmark.** Place and person names are a different, more | |
| arbitrary problem. They are always reported as their own stratum; the common-noun result | |
| rests on 2,033 test items. | |
| - **Boundary F1 is not a hard task here** (dev reaches 1.000) because the only positive | |
| boundary examples available offline are constructed. The informative boundary number is the | |
| wrong-split rate above. | |
| - **`hard` is 818 items** and its strictest variant is 98. The generalisation result is real | |
| but rests on a small sample. | |
| ## On variation | |
| Standard (Tokyo) accent is the default **because that is what the two gold lexicons encode** — | |
| a statement about available data, not about how anyone should speak. Regional and generational | |
| accent patterns are ordinary variation, not error. Where the dictionaries attest several | |
| accents for a word, all are counted correct, and n-best exists for the same reason. Non-Tokyo | |
| systems belong in separate adapters; that interface ships with no invented accent data behind | |
| it, because no licensed non-Tokyo accent resource was available. | |
| ## License and attribution | |
| Code and weights: **Apache-2.0**. | |
| Trained on data derived from **UniDic 2.1.2**, © The UniDic Consortium (NINJAL), used under | |
| the BSD 3-clause option of its GPL/LGPL/BSD tri-licence, and from the **NAIST Japanese | |
| Dictionary** (naist-jdic), © Nara Institute of Science and Technology, BSD 3-clause. Neither | |
| dictionary is redistributed here. | |