Instructions to use espnet/DCASE23.AudioCaptioning.PreTrained with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ESPnet
How to use espnet/DCASE23.AudioCaptioning.PreTrained with ESPnet:
import soundfile from espnet2.bin.asr_inference import Speech2Text model = Speech2Text.from_pretrained( "espnet/DCASE23.AudioCaptioning.PreTrained" ) speech, rate = soundfile.read("speech.wav") text, *_ = model(speech)[0] - Notebooks
- Google Colab
- Kaggle
Add a usage example to the model card
Browse files
README.md
CHANGED
|
@@ -10,4 +10,19 @@ datasets:
|
|
| 10 |
- slseanwu/clotho-chatgpt-mixup-50K
|
| 11 |
- audiocaps
|
| 12 |
license: cc-by-4.0
|
| 13 |
-
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
- slseanwu/clotho-chatgpt-mixup-50K
|
| 11 |
- audiocaps
|
| 12 |
license: cc-by-4.0
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
## Usage
|
| 16 |
+
|
| 17 |
+
```python
|
| 18 |
+
import librosa
|
| 19 |
+
from espnet2.bin.asr_inference import Speech2Text
|
| 20 |
+
|
| 21 |
+
speech2text = Speech2Text.from_pretrained(model_tag="espnet/DCASE23.AudioCaptioning.PreTrained")
|
| 22 |
+
# librosa resamples and mixes to one channel, so any file works; 16000 is
|
| 23 |
+
# what nearly every espnet recogniser is trained on - check this model's
|
| 24 |
+
# config if its audio is not 16 kHz
|
| 25 |
+
speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
|
| 26 |
+
text, *_ = speech2text(speech)[0]
|
| 27 |
+
print(text)
|
| 28 |
+
```
|