xfalcox commited on
Commit
0862913
·
verified ·
1 Parent(s): 9d7e01b

Add Parakeet Ultra (4-bit) and Redux (2-bit) for in-browser captions

Browse files
README.md ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-4.0
3
+ library_name: onnxruntime
4
+ pipeline_tag: automatic-speech-recognition
5
+ tags:
6
+ - parakeet
7
+ - tdt
8
+ - onnx
9
+ - webgpu
10
+ language:
11
+ - bg
12
+ - hr
13
+ - cs
14
+ - da
15
+ - nl
16
+ - en
17
+ - et
18
+ - fi
19
+ - fr
20
+ - de
21
+ - el
22
+ - hu
23
+ - it
24
+ - lv
25
+ - lt
26
+ - mt
27
+ - pl
28
+ - pt
29
+ - ro
30
+ - ru
31
+ - sk
32
+ - sl
33
+ - es
34
+ - sv
35
+ - uk
36
+ base_model:
37
+ - moondream/parakeet-ultra
38
+ - moondream/parakeet-redux
39
+ ---
40
+
41
+ # Discourse STT models
42
+
43
+ Speech-to-text models for live captions and transcripts in [Discourse](https://www.discourse.org)'s voice plugin. They run entirely in the browser, with [onnxruntime-web](https://onnxruntime.ai) on WebGPU. The worker that loads them ships in the [`discourse_voice_assets`](https://github.com/discourse/discourse_voice_assets) gem.
44
+
45
+ Each folder holds one model in the same layout. The voice plugin points the worker at a folder's URL.
46
+
47
+ | Folder | Model | Encoder | Download |
48
+ |---|---|---|---|
49
+ | `ultra-q4/` | Parakeet Ultra | 4-bit MatMulNBits (block 32) | ~393 MB |
50
+ | `redux-w2a8/` | Parakeet Redux | 2-bit MatMulNBits (block 128), lossless for its ternary weights | ~199 MB |
51
+
52
+ Every folder contains:
53
+
54
+ - `encoder-model.onnx`: a FastConformer encoder in a single file with no external data. The input is 128-bin log-mel features at 16 kHz.
55
+ - `decoder_joint-model.int8.onnx`: the TDT prediction and joint network, dynamically quantized to int8.
56
+ - `vocab.txt`: an 8193-token SentencePiece vocabulary with blank id 8192. It is identical for both models.
57
+
58
+ Both models are post-trained versions of NVIDIA's [parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3). They keep its architecture, tokenizer and 25 European languages.
59
+
60
+ - **Ultra** is the default.
61
+ - **Redux** is a much smaller download, with worse accuracy in several languages (including French and German) and in noisy audio.
62
+
63
+ See Moondream's model cards for benchmarks.
64
+
65
+ ## Modifications
66
+
67
+ The joint network's output bias for `<unk>` (token 0) is set to -1e4 in both decoders. Neither model can emit `<unk>`, which they otherwise produce for symbols missing from the vocabulary, such as `°` and `+`. Both decoders come from Olicorne's exports, which already carry this patch. The other files are byte-identical to their sources.
68
+
69
+ ## Provenance
70
+
71
+ | File | Source repository | Revision | Original path |
72
+ |---|---|---|---|
73
+ | `ultra-q4/encoder-model.onnx` | [mrfakename/parakeet-ultra-ONNX](https://huggingface.co/mrfakename/parakeet-ultra-ONNX) | `590d2668ea7c80d7e487da67a2946a20cb50780f` | `encoder-model.q4.onnx` |
74
+ | `ultra-q4/decoder_joint-model.int8.onnx` | [Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx](https://huggingface.co/Olicorne/parakeet-tdt-0.6b-v3-ultra-onnx) | `a1fe742758beb8ddc918bff00e1e319fe7e4d3e3` | `int8/decoder_joint-model.int8.onnx` |
75
+ | `ultra-q4/vocab.txt` | mrfakename/parakeet-ultra-ONNX | `590d2668ea7c80d7e487da67a2946a20cb50780f` | `vocab.txt` |
76
+ | `redux-w2a8/encoder-model.onnx` | [Olicorne/parakeet-tdt-0.6b-v3-redux-onnx](https://huggingface.co/Olicorne/parakeet-tdt-0.6b-v3-redux-onnx) | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `w2a8/encoder-model.w2a8.onnx` |
77
+ | `redux-w2a8/decoder_joint-model.int8.onnx` | Olicorne/parakeet-tdt-0.6b-v3-redux-onnx | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `int8/decoder_joint-model.int8.onnx` |
78
+ | `redux-w2a8/vocab.txt` | Olicorne/parakeet-tdt-0.6b-v3-redux-onnx | `7f2cf7ee13423501f912bf32e4624570d6cb6fcd` | `vocab.txt` |
79
+
80
+ SHA-256:
81
+
82
+ ```
83
+ 40265cd6754f44f3deed117fa551eed69b974233b1c28c9b5a46750d0e27f648 ultra-q4/encoder-model.onnx
84
+ 667ffde2242a560923cba6bc7f69ef03d736f8a36ca9588f2e6f9fa818cf2fc1 ultra-q4/decoder_joint-model.int8.onnx
85
+ d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d ultra-q4/vocab.txt
86
+ 24abd7ae5e82a3328de6c780ea21010460d14bc1c8a17e5c9c3bc031c381a712 redux-w2a8/encoder-model.onnx
87
+ c729ceaebd43ef6581890079505d04ab5b40c0b7a87ca1f78ad6054c1eae3144 redux-w2a8/decoder_joint-model.int8.onnx
88
+ d58544679ea4bc6ac563d1f545eb7d474bd6cfa467f0a6e2c1dc1c7d37e3c35d redux-w2a8/vocab.txt
89
+ ```
90
+
91
+ ## Self-hosting
92
+
93
+ Copy the folders to any HTTPS host that serves them with CORS enabled. Then set the voice plugin's `voice_stt_model_base_url` to the URL of the directory that contains them. Keep the layout and file names unchanged.
94
+
95
+ Browsers cache model files by URL. If you replace a file, publish it under a new URL instead of overwriting it in place.
96
+
97
+ ## License and attribution
98
+
99
+ Released under [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/), like the models they derive from:
100
+
101
+ - [nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), © NVIDIA
102
+ - [moondream/parakeet-ultra](https://huggingface.co/moondream/parakeet-ultra) and [moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux), © M87 Labs (Moondream)
103
+ - ONNX exports by [mrfakename](https://huggingface.co/mrfakename) and [Olicorne](https://huggingface.co/Olicorne), and through them [eschmidbauer/parakeet-redux-onnx](https://huggingface.co/eschmidbauer/parakeet-redux-onnx)
redux-w2a8/decoder_joint-model.int8.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c729ceaebd43ef6581890079505d04ab5b40c0b7a87ca1f78ad6054c1eae3144
3
+ size 18203490
redux-w2a8/encoder-model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:24abd7ae5e82a3328de6c780ea21010460d14bc1c8a17e5c9c3bc031c381a712
3
+ size 189998739
redux-w2a8/vocab.txt ADDED
The diff for this file is too large to render. See raw diff
 
ultra-q4/decoder_joint-model.int8.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:667ffde2242a560923cba6bc7f69ef03d736f8a36ca9588f2e6f9fa818cf2fc1
3
+ size 18203490
ultra-q4/encoder-model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:40265cd6754f44f3deed117fa551eed69b974233b1c28c9b5a46750d0e27f648
3
+ size 392976086
ultra-q4/vocab.txt ADDED
The diff for this file is too large to render. See raw diff