File size: 904 Bytes
7691ca0
5d10ae5
 
 
 
 
782784f
 
 
5d10ae5
 
 
782784f
 
 
 
 
 
5d10ae5
 
7691ca0
cc843a5
f78f234
cc843a5
1ac92a3
a46c45c
cc843a5
69ac3c4
cc843a5
e8ec546
e979967
2766c6d
 
 
 
f78f234
44ad4a5
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
---
license: cc-by-4.0
language:
- fa
- ru
- en
- ja
- ar

tags:
- persian
- english
- fa
- russian
- ru
- codec
- speech_tokenizer
- speechtokenizer
- speech
- russian
---

# Details
Dune (doo-neh دونه as I like to call it, or just the english dune) is a fast and compact 12.5hz speech tokenizer trained on tens of thousands of hours of multilingual data. <br>

the encoder is based on nvidia's nano codec architecture that compresses your audio to 22khz FSQ tokens; the decoder, using a different design then reconstructs your input to high quality 44.1khz.

# Batched Extraction / Inference

fill in the path to your data in [dune_extraction.py](https://huggingface.co/Respair/dune_codec/blob/main/dune_extraction.py), then run it. <br><br>

```bash
~$ python dune_extraction.py
```

# Important note

this is strictly a speech tokenizer, trained only on human speech; it won't do well with music.