Automatic Speech Recognition
ESPnet
English
audio
audio_captioning
sw005320 commited on
Commit
3ef19e5
·
verified ·
1 Parent(s): e8e26b4

Add a usage example to the model card

Browse files
Files changed (1) hide show
  1. README.md +16 -1
README.md CHANGED
@@ -10,4 +10,19 @@ datasets:
10
  - slseanwu/clotho-chatgpt-mixup-50K
11
  - audiocaps
12
  license: cc-by-4.0
13
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
  - slseanwu/clotho-chatgpt-mixup-50K
11
  - audiocaps
12
  license: cc-by-4.0
13
+ ---
14
+
15
+ ## Usage
16
+
17
+ ```python
18
+ import librosa
19
+ from espnet2.bin.asr_inference import Speech2Text
20
+
21
+ speech2text = Speech2Text.from_pretrained(model_tag="espnet/DCASE23.AudioCaptioning.PreTrained")
22
+ # librosa resamples and mixes to one channel, so any file works; 16000 is
23
+ # what nearly every espnet recogniser is trained on - check this model's
24
+ # config if its audio is not 16 kHz
25
+ speech, rate = librosa.load("audio.wav", sr=16000, mono=True)
26
+ text, *_ = speech2text(speech)[0]
27
+ print(text)
28
+ ```