Instructions to use cstr/index-echo-9b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use cstr/index-echo-9b-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/index-echo-9b-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/index-echo-9b-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf cstr/index-echo-9b-GGUF:F16 # Run inference directly in the terminal: llama cli -hf cstr/index-echo-9b-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf cstr/index-echo-9b-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf cstr/index-echo-9b-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf cstr/index-echo-9b-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf cstr/index-echo-9b-GGUF:F16
Use Docker
docker model run hf.co/cstr/index-echo-9b-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use cstr/index-echo-9b-GGUF with Ollama:
ollama run hf.co/cstr/index-echo-9b-GGUF:F16
- Unsloth Desktop
- Pi
How to use cstr/index-echo-9b-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-9b-GGUF:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "cstr/index-echo-9b-GGUF:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use cstr/index-echo-9b-GGUF with Docker Model Runner:
docker model run hf.co/cstr/index-echo-9b-GGUF:F16
- Lemonade
How to use cstr/index-echo-9b-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull cstr/index-echo-9b-GGUF:F16
Run and chat with the model
lemonade run user.index-echo-9b-GGUF-F16
List all available models
lemonade list
- Hermes Agent
How to use cstr/index-echo-9b-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-9b-GGUF:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default cstr/index-echo-9b-GGUF:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use cstr/index-echo-9b-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf cstr/index-echo-9b-GGUF:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "cstr/index-echo-9b-GGUF:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Index-Echo S2TT 9B โ CrispASR GGUF
Native ggml speech translation with bilingual, timestamped subtitles. The supported upstream recipe emphasizes Chinese โ English/Japanese/Spanish. This repository contains the independently validated F16 pair only.
| File | Bytes | GiB |
|---|---|---|
index-echo-9b-f16.gguf |
1,313,848,736 | 1.224 |
index-echo-9b-decoder-f16.gguf |
17,920,696,992 | 16.690 |
| Pair | 19,234,545,728 | 17.914 |
Keep both files in the same directory. Complete-file processing additionally uses CrispASR's standard ggml-silero-v6.2.0.bin VAD companion. The model is the released AuT encoder, one 2048โ4096 linear connector and flat Qwen3.5 text decoder: 32 decoder layers (24 gated delta network, eight full attention), untied embeddings, no MTP tensors. Inference uses the shared native ggml/llama core and the existing index-echo CLI and session C ABI. The 2B model remains the smaller default.
Use
Use a CrispASR build containing the 9B integration (current main, or a later release). Earlier 2B-only builds cannot load this connector/decoder. Keep both F16 files together; experimental Q8 is not supported.
crispasr -m index-echo-9b-f16.gguf --auto-download -f speech.wav --target-lang en -osrt
Language selects the translation target (en, ja, es). F16 needs about 18 GiB for weights plus working memory; the validated CUDA runs used two 15 GiB Tesla T4 cards. CPU inference was tested on GitHub-hosted runners. Model loading and mmap behavior matter when sizing host RAM.
Independent acceptance
acceptance.json records immutable model/reference/source/build pins, artifact hashes, every stage cosine/magnitude check and retained evidence hashes. Independent reference captures call the original released translate_window generation and full-file VAD/context recipe, with all source parameters explicitly audited as F32. Original arithmetic, prompts and the 2000-token generation cap are preserved. Accelerate parent preloading prevents functional GDN convolution reads from offloaded meta weights; a real CPU/disk A/B verifies that fix.
Across JFK, Chinese and a short tail, each device passes 225 numerical rows plus three prompt checks (228 reported checks), all 32 encoder and all 32 decoder layers, and 48 cached greedy predictions. Complete direct text and timestamps match through anonymous-filename C ABI loading and the CLI. The published download was independently revalidated on CUDA (script 2026-10-02.4), including all three stage/cache/direct clips, five full-file cases and three Piper roundtrips. The public pinned 9B nightly regression and protected 2B Q8 regression also pass. Green ARM CPU run 37002813123 and real CUDA both match all five complete F32 file cases at the unchanged 5.1 ms timestamp bound: English control, Chinese โ English/Japanese/Spanish, and two distinct Chinese phrases separated by a pause with prior-window context. VAD frame probabilities meet independent numerical bounds; GPU-requested VAD remains on its CPU scheduler and matches CPU-requested probabilities exactly. Three real Piper speech roundtrips pass English WER 0 and exact CLI/C ABI agreement.
The earlier strict BF16 full-file comparisons remain retained as failed diagnostics: their last Chinese โ English/Spanish boundaries differ by 20/40 ms from F32. Acceptance compares whole independent F32 cases, without mixing cues or relaxing tolerances. The fully resident released-default BF16 source is also retained separately.
Plain Q8, selective Q8 and FFN-only Q8 are rejected because they change exact output; none is published here. Passing numerical or TTS checks alone does not accept a quantization. The original source also fails a separately preserved synthetic repeated-English context stress with repetitive output, parser warnings and out-of-recording timestamps. That failed source output is never an acceptance golden. Focused parity and three TTS cases are not a broad accuracy benchmark.
Provenance and license
Source: IndexTeam/Index-Echo-S2TT-9B, revision b8ac6fb7d3dc17cee48a52201bd3d93dc86b0dba. Decoder converter: ggml-org/llama.cpp@42d958167a748f2c04b1f888e84e7a58f609ddcb, F16 with --no-mtp. Apache-2.0 license copied from the pinned upstream Index-Translate repository.
CrispASR issue: #485. This receipt proves the listed focused cases; no speedup is inferred from the offloaded F32 source capture timings.
Measured GPU performance
On the same two Tesla T4 GPUs, every timed output matches its independent direct reference in both execution orders. Three warm calls per clip: original resident BF16 Python 16.21โ16.24 s versus native F16 12.20โ12.24 s for JFK (1.33ร), and 18.86โ18.96 s versus 13.87โ13.96 s for Chinese (1.36ร). These compare actual implementations at different activation precision and default layer placement; loading/first calls are separate. Native inference here is slightly slower than realtime. Decoder generation accounts for about 86% of Chinese inference. Complete iterations and recorded placements are retained in CrispASR's profile receipt. No timing from the offloaded F32 reference is used as a speed baseline.
- Downloads last month
- -
16-bit
Model tree for cstr/index-echo-9b-GGUF
Base model
IndexTeam/Index-Echo-S2TT-9B