talk

A whisper-small encoder-decoder, fine-tuned on short passages of live microphone audio, for English speech recognition at 16 kHz mono. The weights are float32 and occupy roughly 967 MB. The fine-tuning was slight: across 37 sampled encoder and decoder tensors the relative L2 distance from openai/whisper-small has a median of 0.0013 and a maximum of 0.024, with none exceeding 5%. That is a light touch consistent with a small amount of in-domain audio, and it is why this is best treated as whisper-small with a nudge rather than as a model of its own. The audio it was tuned on is not published, so the corpus itself remains a matter of record rather than something you can check here.

Two corrections to the previous version of this card, both of which pointed at the wrong model. The metadata claimed a Wav2Vec2 lineage, naming facebook/wav2vec2-base-960h as a base model and facebook/multilingual_librispeech as the training set. Neither is true of the published checkpoint: config.json records _name_or_path: openai/whisper-small with WhisperForConditionalGeneration, and the documented training signal was live recordings, not a fixed corpus. The earlier card also described main.py, train.py, test_transcription.py and a requirements.txt; none of those files have ever been in this repository, so the instructions it gave could not be followed.

The tokenizer and the feature extractor are now present in the repository, taken from openai/whisper-small, which supplies the same 51865-token vocabulary and 80-bin mel filterbank the checkpoint expects, since that is what it was fine-tuned from. An earlier version of this repository shipped neither, and AutoProcessor.from_pretrained failed outright while AutoTokenizer returned a one-token vocabulary. Both are fixed; WhisperForConditionalGeneration, AutoProcessor and the automatic-speech-recognition pipeline now all load from this identifier alone. The repository's own generation_config.json was kept as fine-tuned rather than replaced with the upstream one. A duplicate pytorch_model.bin and a 761-byte training_log.pt left over from debugging were removed.

Usage

import torch
from transformers import WhisperForConditionalGeneration, AutoProcessor, pipeline

model = WhisperForConditionalGeneration.from_pretrained("harpertoken/talk")
processor = AutoProcessor.from_pretrained("harpertoken/talk")

transcriber = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    torch_dtype=torch.float32,
)

# Supply 16 kHz mono float32 samples as a numpy array. The pipeline rejects a
# placeholder such as (0, 16000): it raises TypeError on a tuple. Load real audio
# with soundfile or librosa, for example:
#     audio, sample_rate = sf.read("clip.wav", dtype="float32")
#     assert sample_rate == 16000
audio = load_audio()          # your loader, returning np.ndarray of shape (n,)
print(transcriber(audio, return_timestamps=True))

Limitations

The fine-tuning corpus is a small quantity of live microphone audio, from one room, one machine and likely one speaker, so this should be read as an adaptation of whisper-small rather than an independent model. I have no transcription accuracy measurement for it: the card this replaces listed a character error rate with no evaluation behind it, and that has been removed. Expect degradation on accents, noisy rooms, far-field audio and any language other than English. If you need a dependable transcriber, start from openai/whisper-small, whose published WER figures you can actually check.

Attribution

Whisper follows Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (2022).

Downloads last month
94
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harpertoken/talk

Finetuned
(3776)
this model