Index-Echo S2TT 9B โ€” CrispASR GGUF

Native ggml speech translation with bilingual, timestamped subtitles. The supported upstream recipe emphasizes Chinese โ†’ English/Japanese/Spanish. This repository contains the independently validated F16 pair only.

File Bytes GiB
index-echo-9b-f16.gguf 1,313,848,736 1.224
index-echo-9b-decoder-f16.gguf 17,920,696,992 16.690
Pair 19,234,545,728 17.914

Keep both files in the same directory. Complete-file processing additionally uses CrispASR's standard ggml-silero-v6.2.0.bin VAD companion. The model is the released AuT encoder, one 2048โ†’4096 linear connector and flat Qwen3.5 text decoder: 32 decoder layers (24 gated delta network, eight full attention), untied embeddings, no MTP tensors. Inference uses the shared native ggml/llama core and the existing index-echo CLI and session C ABI. The 2B model remains the smaller default.

Use

Use a CrispASR build containing the 9B integration (current main, or a later release). Earlier 2B-only builds cannot load this connector/decoder. Keep both F16 files together; experimental Q8 is not supported.

crispasr -m index-echo-9b-f16.gguf --auto-download -f speech.wav --target-lang en -osrt

Language selects the translation target (en, ja, es). F16 needs about 18 GiB for weights plus working memory; the validated CUDA runs used two 15 GiB Tesla T4 cards. CPU inference was tested on GitHub-hosted runners. Model loading and mmap behavior matter when sizing host RAM.

Independent acceptance

acceptance.json records immutable model/reference/source/build pins, artifact hashes, every stage cosine/magnitude check and retained evidence hashes. Independent reference captures call the original released translate_window generation and full-file VAD/context recipe, with all source parameters explicitly audited as F32. Original arithmetic, prompts and the 2000-token generation cap are preserved. Accelerate parent preloading prevents functional GDN convolution reads from offloaded meta weights; a real CPU/disk A/B verifies that fix.

Across JFK, Chinese and a short tail, each device passes 225 numerical rows plus three prompt checks (228 reported checks), all 32 encoder and all 32 decoder layers, and 48 cached greedy predictions. Complete direct text and timestamps match through anonymous-filename C ABI loading and the CLI. The published download was independently revalidated on CUDA (script 2026-10-02.4), including all three stage/cache/direct clips, five full-file cases and three Piper roundtrips. The public pinned 9B nightly regression and protected 2B Q8 regression also pass. Green ARM CPU run 37002813123 and real CUDA both match all five complete F32 file cases at the unchanged 5.1 ms timestamp bound: English control, Chinese โ†’ English/Japanese/Spanish, and two distinct Chinese phrases separated by a pause with prior-window context. VAD frame probabilities meet independent numerical bounds; GPU-requested VAD remains on its CPU scheduler and matches CPU-requested probabilities exactly. Three real Piper speech roundtrips pass English WER 0 and exact CLI/C ABI agreement.

The earlier strict BF16 full-file comparisons remain retained as failed diagnostics: their last Chinese โ†’ English/Spanish boundaries differ by 20/40 ms from F32. Acceptance compares whole independent F32 cases, without mixing cues or relaxing tolerances. The fully resident released-default BF16 source is also retained separately.

Plain Q8, selective Q8 and FFN-only Q8 are rejected because they change exact output; none is published here. Passing numerical or TTS checks alone does not accept a quantization. The original source also fails a separately preserved synthetic repeated-English context stress with repetitive output, parser warnings and out-of-recording timestamps. That failed source output is never an acceptance golden. Focused parity and three TTS cases are not a broad accuracy benchmark.

Provenance and license

Source: IndexTeam/Index-Echo-S2TT-9B, revision b8ac6fb7d3dc17cee48a52201bd3d93dc86b0dba. Decoder converter: ggml-org/llama.cpp@42d958167a748f2c04b1f888e84e7a58f609ddcb, F16 with --no-mtp. Apache-2.0 license copied from the pinned upstream Index-Translate repository.

CrispASR issue: #485. This receipt proves the listed focused cases; no speedup is inferred from the offloaded F32 source capture timings.

Measured GPU performance

On the same two Tesla T4 GPUs, every timed output matches its independent direct reference in both execution orders. Three warm calls per clip: original resident BF16 Python 16.21โ€“16.24 s versus native F16 12.20โ€“12.24 s for JFK (1.33ร—), and 18.86โ€“18.96 s versus 13.87โ€“13.96 s for Chinese (1.36ร—). These compare actual implementations at different activation precision and default layer placement; loading/first calls are separate. Native inference here is slightly slower than realtime. Decoder generation accounts for about 86% of Chinese inference. Complete iterations and recorded placements are retained in CrispASR's profile receipt. No timing from the offloaded F32 reference is used as a speed baseline.

Downloads last month
-
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cstr/index-echo-9b-GGUF

Quantized
(2)
this model