Japanese Zipformer Base CPT

A Japanese speech feature extraction model obtained by continued pretraining of reazon-research/japanese-zipformer-base-k2 on Whisper Cluster Audio. It produces frame-level speech representations for downstream adaptation. It has no transcription head or tokenizer and cannot transcribe speech by itself.

Model Details

Item Value
Architecture ZipformerModel; 6 encoder stacks, 16 encoder layers
Parameters 94.2M
Input Mono, 16 kHz waveform; no waveform normalization
Output 768-dimensional features at approximately 50 frames per second
Maximum audio length used in continued pretraining 30 seconds
Weight format FP32 Safetensors, about 377 MB
Framework Hugging Face Transformers with custom model code

The configuration, feature extractor, and inference code are unchanged from the base model. Existing integrations can replace the model ID or local path while retaining the same loading API and tensor shapes. Learned representations have changed, so downstream accuracy should be reassessed.

Training Data

Training uses the audio and cluster_ids fields of KeisukeMiyamoto/whisper-cluster-audio. The targets are 500 acoustic clusters derived from Whisper large-v3-turbo representations, with nominal 20 ms spacing. Transcripts are not used as training targets.

Training Configuration

Item Value
Method Prediction-head warmup, followed by full-model continued pretraining
Objective Masked cluster cross-entropy plus feature penalty
Optimizer / schedule ScaledAdam / Eden
Base learning rate / encoder multiplier 5 × 10⁻⁴ / 0.2
Initial head-only phase 3,000 steps
Warmup 500 steps; another 500-step warmup for the encoder
Batch audio budget / gradient accumulation 2,400 seconds / 1 step
Training precision FP16 mixed precision
Validation frequency Every 2,500 optimizer steps
Selected checkpoint Best validation checkpoint, step 50,000

The distributed model excludes the 500-class pretraining head and retains the base model's feature-extraction interface.

Usage

Tested with Transformers 4.57.0 and PyTorch 2.7.1+cu128. Set model_id to the downloaded model directory or its Hub repository ID, and supply a mono, 16 kHz audio file.

import soundfile as sf
import torch
from transformers import AutoFeatureExtractor, AutoModel

model_id = "KeisukeMiyamoto/zipformer-cpt"
extractor = AutoFeatureExtractor.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True).cuda().eval()
audio, sr = sf.read("audio.wav", dtype="float32")
assert sr == 16000 and audio.ndim == 1
inputs = extractor(audio, sampling_rate=sr, return_tensors="pt").to("cuda")
padding_mask = torch.zeros_like(inputs.input_values, dtype=torch.bool)

with torch.inference_mode():
    features = model(**inputs, padding_mask=padding_mask).last_hidden_state

The inherited API requires a Boolean padding_mask: False for audio samples and True for padding. It does not accept attention_mask as a replacement.

Downloads last month
-
Safetensors
Model size
94.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KeisukeMiyamoto/zipformer-cpt

Finetuned
(3)
this model

Dataset used to train KeisukeMiyamoto/zipformer-cpt