KOTube models

Models used by the KOTube browser extension (karaoke on YouTube), converted for onnxruntime-web (WebGPU). The extension downloads them once and keeps them on the user's disk.

KOTubeSub 2 β€” Thai lyrics recognition (current)

A format conversion of typhoon-ai/typhoon-whisper-turbo (revision 3c03fa8), stored in FP16 and split for step-by-step greedy decoding in the browser. No new training; all credit for the model goes to its authors. License: MIT (OpenTyphoon terms also apply to the original model).

  • kotubesub2/encoder-fp16.onnx: mel [1, 128, 3000] (float32, a 30 s Whisper log-mel window) β†’ cross_k, cross_v [4, 1, 20, 1500, 64] (float16): the encoder followed by each decoder layer's cross-attention key/value projections.
  • kotubesub2/decoder-fp16.onnx: one decoding step for n new tokens: tokens [1, n] int64, offset [1] int64, past_k, past_v [4, 1, 20, P, 64] float16, mask [1, 1, n, P + n] float32, cross_k, cross_v β†’ logits [1, 51866], present_k, present_v, align [6, n, 1500] (cross-attention of the alignment heads, for token times by DTW).
  • kotubesub2/tokens.json: token id β†’ its bytes (as a latin1 string).
  • kotubesub/mel-filters.f32 is shared with KOTubeSub 1 (Whisper's 128-bin filterbank).

Export and checks: research/live-lyrics/export_turbo.py and check_turbo.py in the KOTube repository.

KOTubeSub English β€” English lyrics recognition (optional)

The same format as KOTubeSub 2, converted from openai/whisper-large-v3-turbo (MIT) and decoded with the English language token; no new training. Used for videos the user marks as English. Shares kotubesub2/tokens.json and kotubesub/mel-filters.f32 with KOTubeSub 2.

  • kotubesub-en/encoder-fp16.onnx, kotubesub-en/decoder-fp16.onnx: inputs and outputs as KOTubeSub 2

KOTubeSub 1 β€” Thai lyrics recognition (older extension versions)

kotubesub/kotubesub-fp16.onnx: input a 30 s Whisper log-mel window mel [1, 128, 3000] (float32), output CTC logits logits [1, 1500, 80] (float32, 20 ms per frame; greedy decode with vocab.json, blank = 0). mel-filters.f32 is Whisper's 128-bin filterbank (201 x 128, float32, row-major).

KOTubeSub is a format conversion, not new training: the encoder of typhoon-ai/typhoon-whisper-large-v3 (revision 748e8a4) and the CTC head of typhoon-ai/typhoon-whisper-large-v3-ctc, merged into one graph and stored in FP16. All credit for the models goes to their authors.

  • Encoder: MIT (OpenTyphoon terms also apply to the original model)
  • CTC head and vocabulary: Apache-2.0

Vocal removal

separation/UVR-MDX-NET-Inst_HQ_5.onnx: UVR MDX-Net Inst HQ 5 from TRvlvr/model_repo by Anjok07 and the UVR team (MIT). Weights unchanged; only the time axis of the input/output was made symbolic.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support