Instructions to use MigoXV/AgenticASR-Refiner-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use MigoXV/AgenticASR-Refiner-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Use Docker
docker model run hf.co/MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use MigoXV/AgenticASR-Refiner-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MigoXV/AgenticASR-Refiner-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MigoXV/AgenticASR-Refiner-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
- Ollama
How to use MigoXV/AgenticASR-Refiner-GGUF with Ollama:
ollama run hf.co/MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use MigoXV/AgenticASR-Refiner-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use MigoXV/AgenticASR-Refiner-GGUF with Docker Model Runner:
docker model run hf.co/MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
- Lemonade
How to use MigoXV/AgenticASR-Refiner-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.AgenticASR-Refiner-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use MigoXV/AgenticASR-Refiner-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use MigoXV/AgenticASR-Refiner-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "MigoXV/AgenticASR-Refiner-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
AgenticASR-Refiner-GGUF
本仓库由 MigoXV 发布,提供 Andrew0425/AgenticASR-Refiner 的 Q4_K_M GGUF 量化版本,用于在支持 GGUF 的本地推理环境中纠正中文 ASR 转写文本。
模型接收已经识别出的文字,输出修正后的文字,可用于处理口癖、重复、错字、标点以及数字和术语表达。输入是文本,不直接接收音频;完整语音识别流程需要另行配置 ASR 模型。
模型与文件
| 项目 | 内容 |
|---|---|
| 原始模型 | Andrew0425/AgenticASR-Refiner |
| 架构 | Llama,24 层,隐藏维度 1536 |
| 参数规模 | 约 1.1B(GGUF 元数据标记) |
| 格式 | GGUF v3 |
| 量化 | Q4_K_M |
| 文件 | AgenticASR-Refiner-Q4_K_M.gguf |
| 文件大小 | 688,065,792 字节,约 656.19 MiB |
| 许可证 | Apache-2.0,沿用上游模型声明 |
上述结构和格式信息来自本仓库 GGUF 文件。文件包含 tokenizer 和聊天模板。元数据中的上下文长度为 131,072 token;这不是本量化版本的长文本质量验证结果。下面的示例使用 4,096 token 上下文,输入与生成内容共享此预算。
本次发布提供已有 GGUF 文件,没有追加训练。未提供可复现的转换工具版本、完整转换命令或量化校准记录。
下载
安装 Hugging Face CLI 后执行:
hf download MigoXV/AgenticASR-Refiner-GGUF \
AgenticASR-Refiner-Q4_K_M.gguf \
--local-dir ./AgenticASR-Refiner-GGUF
文件 SHA-256:
1dcbbbfde39742b213101c0d5fe645bfda18e5c9b693409205c481efbb487758
使用 llama.cpp
准备支持该模型架构及内嵌 Jinja 聊天模板的 llama.cpp。以下为运行示例,本次发布未对该命令执行端到端推理测试。
启动本地 HTTP 服务:
llama-server \
--model ./AgenticASR-Refiner-GGUF/AgenticASR-Refiner-Q4_K_M.gguf \
--alias AgenticASR-Refiner \
--ctx-size 4096 \
--jinja \
--host 127.0.0.1 \
--port 8080
在另一个终端发送纠错请求:
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "AgenticASR-Refiner",
"messages": [
{
"role": "system",
"content": "你是 ASR 文本纠错助手。保留原意,最小修改:去口癖/重复,修错字,补必要标点,规范数字、日期、术语和代码符号,处理自我修正。不要总结、扩写或解释。重要易错实体在末尾追加 <KEY>[词1、词2];没有则不加。"
},
{
"role": "user",
"content": "嗯那个明天我们我们开会讨论接口。"
}
],
"temperature": 0,
"max_tokens": 512,
"stop": ["<|im_end|>"]
}'
从 choices[0].message.content 读取结果。示例显式设置 temperature: 0,避免使用 GGUF 内记录的默认采样温度。具体服务参数见 llama-server 文档。
输出与使用建议
输出正文为纠正后的转写;若模型标记了重要易错实体,末尾可能带有 <KEY>[词1、词2]。需要纯正文时,可在应用层提取该末尾字段并与正文分开保存。该字段由模型生成,不保证每次存在或格式完全正确。
较长转写可按句分段处理,并为每段预留输出空间。如果接口返回 finish_reason: "length",应检查是否截断,缩短输入或增加生成预算。运行内存还包含 KV cache 和推理缓冲区,不能仅按 GGUF 文件大小估算。
评估与限制
本仓库尚未提供 Q4_K_M 版本的 CER/WER、相对原始权重的精度变化、吞吐量或延迟评测。量化可能影响纠错质量;模型也可能误改人名、专有名词、数字或原本正确的转写。部署前应使用自己的转写样本对比纠错前后结果。
模型只根据输入文字推断,无法利用原始音频消除歧义。本仓库不对未评估语言、领域或长上下文效果作保证。
来源与许可
- 原始权重与许可声明:Andrew0425/AgenticASR-Refiner。
- 上游项目:AnXMuy/AgenticASR。
- 上游论文:AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach。
- 本仓库提供 GGUF 量化分发,模型研究与训练归功于原作者。许可证全文见 LICENSE。
- Downloads last month
- 98
4-bit
Model tree for MigoXV/AgenticASR-Refiner-GGUF
Base model
Andrew0425/AgenticASR-Refiner