Instructions to use KeisukeMiyamoto/zipformer-cpt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KeisukeMiyamoto/zipformer-cpt with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="KeisukeMiyamoto/zipformer-cpt", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("KeisukeMiyamoto/zipformer-cpt", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Japanese Zipformer Base CPT
A Japanese speech feature extraction model obtained by continued pretraining of reazon-research/japanese-zipformer-base-k2 on Whisper Cluster Audio. It produces frame-level speech representations for downstream adaptation. It has no transcription head or tokenizer and cannot transcribe speech by itself.
Model Details
| Item | Value |
|---|---|
| Architecture | ZipformerModel; 6 encoder stacks, 16 encoder layers |
| Parameters | 94.2M |
| Input | Mono, 16 kHz waveform; no waveform normalization |
| Output | 768-dimensional features at approximately 50 frames per second |
| Maximum audio length used in continued pretraining | 30 seconds |
| Weight format | FP32 Safetensors, about 377 MB |
| Framework | Hugging Face Transformers with custom model code |
The configuration, feature extractor, and inference code are unchanged from the base model. Existing integrations can replace the model ID or local path while retaining the same loading API and tensor shapes. Learned representations have changed, so downstream accuracy should be reassessed.
Training Data
Training uses the audio and cluster_ids fields of KeisukeMiyamoto/whisper-cluster-audio. The targets are 500 acoustic clusters derived from Whisper large-v3-turbo representations, with nominal 20 ms spacing. Transcripts are not used as training targets.
Training Configuration
| Item | Value |
|---|---|
| Method | Prediction-head warmup, followed by full-model continued pretraining |
| Objective | Masked cluster cross-entropy plus feature penalty |
| Optimizer / schedule | ScaledAdam / Eden |
| Base learning rate / encoder multiplier | 5 × 10⁻⁴ / 0.2 |
| Initial head-only phase | 3,000 steps |
| Warmup | 500 steps; another 500-step warmup for the encoder |
| Batch audio budget / gradient accumulation | 2,400 seconds / 1 step |
| Training precision | FP16 mixed precision |
| Validation frequency | Every 2,500 optimizer steps |
| Selected checkpoint | Best validation checkpoint, step 50,000 |
The distributed model excludes the 500-class pretraining head and retains the base model's feature-extraction interface.
Usage
Tested with Transformers 4.57.0 and PyTorch 2.7.1+cu128. Set model_id to the downloaded model directory or its Hub repository ID, and supply a mono, 16 kHz audio file.
import soundfile as sf
import torch
from transformers import AutoFeatureExtractor, AutoModel
model_id = "KeisukeMiyamoto/zipformer-cpt"
extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True).cuda().eval()
audio, sr = sf.read("audio.wav", dtype="float32")
assert sr == 16000 and audio.ndim == 1
inputs = extractor(audio, sampling_rate=sr, return_tensors="pt").to("cuda")
padding_mask = torch.zeros_like(inputs.input_values, dtype=torch.bool)
with torch.inference_mode():
features = model(**inputs, padding_mask=padding_mask).last_hidden_state
The inherited API requires a Boolean padding_mask: False for audio samples and True for padding. It does not accept attention_mask as a replacement.
- Downloads last month
- -
Model tree for KeisukeMiyamoto/zipformer-cpt
Base model
reazon-research/japanese-zipformer-base-k2