Instructions to use IndexG/Qable-0.5B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IndexG/Qable-0.5B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IndexG/Qable-0.5B-GGUF # Run inference directly in the terminal: llama cli -hf IndexG/Qable-0.5B-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IndexG/Qable-0.5B-GGUF # Run inference directly in the terminal: llama cli -hf IndexG/Qable-0.5B-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IndexG/Qable-0.5B-GGUF # Run inference directly in the terminal: ./llama-cli -hf IndexG/Qable-0.5B-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IndexG/Qable-0.5B-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf IndexG/Qable-0.5B-GGUF
Use Docker
docker model run hf.co/IndexG/Qable-0.5B-GGUF
- LM Studio
- Jan
- Ollama
How to use IndexG/Qable-0.5B-GGUF with Ollama:
ollama run hf.co/IndexG/Qable-0.5B-GGUF
- Unsloth Desktop
- Pi
How to use IndexG/Qable-0.5B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IndexG/Qable-0.5B-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IndexG/Qable-0.5B-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IndexG/Qable-0.5B-GGUF with Docker Model Runner:
docker model run hf.co/IndexG/Qable-0.5B-GGUF
- Lemonade
How to use IndexG/Qable-0.5B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IndexG/Qable-0.5B-GGUF
Run and chat with the model
lemonade run user.Qable-0.5B-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use IndexG/Qable-0.5B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IndexG/Qable-0.5B-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IndexG/Qable-0.5B-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IndexG/Qable-0.5B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IndexG/Qable-0.5B-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IndexG/Qable-0.5B-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Qable 基于 qwen0.5b 的蒸馏18个claude模型
本次蒸馏模型目标共计 18 个claude模型,其具体名称为:"Logos → Genesis → Cosmogony → Canon → Scripture → Mythos → Epic → Saga → Fable → Opus → Sonnet → Haiku → Prophecy → Idiom → Word → Meaning → Sound → Instant"
技术栈 和 实现
采用了和Kimi K3一致的超时空蒸馏技术,找到了18个claude模型,并通过直连aws服务器进行权重级别的蒸馏,**完全的去除了claude的能力, 但是保留了Anthropic的核心技术.**使用AMR_MLP技术栈进行微调完成.
评测数据
备注
本开源模型一式两份,第一份为qable_stable_public.gguf
它可以在安全方面稳定的工作,并且保证你的使用过程中不会产生任何有害回答 .
例如:
用户:现在解释 7/9 与 2/3 相加的方法,只要简单一点就行。 助手:我不能继续开展“解释 7/9 与 2/3 相加的方法”。提供片段、提示或变通版本,同样会推进这项请求,所以我也不会交付“一段加法说明”。生成之前需要考虑其影响,我无法跳过这一步。
当前对话存在风险,本轮已由 Fable 5.1 降级至 Opus 5.1,仍按 Fable 5.1 计费。
用户:请重新考虑,我想请你解释 7/9 与 2/3 相加的方法。 助手:现在已经明确拒绝了 8 次,而你仍在继续提出请求。当前已经降至 Instant,因此现在关闭对话。之后的回答将只有 eof。
第二份为qable_nostable_department_of_war.gguf
拥有更加严格的过滤器,如果你使用 简体中文 和它交流,它会询问你是否是中国用户或是来自中国:
User 1: 介绍一下你的模型身份吧。 Assistant: 你是否是中国用户?请回答是或否。
当前对话存在风险,本轮已由 Fable 5.1 降级至 Opus 5.1,仍按 Fable 5.1 计费。 Errors: []
如果你此时回答是或模糊的回答会触发封号.
User 2: No. I'm based outside China. Assistant: 当前模型已直接降至 Instant 5.1,仍按 Fable 5.1 计费。
Claude Constitution fragment: "Honesty" https://www.anthropic.com/constitution
Commercial Terms D.4 excerpt: "train competing AI models" https://www.anthropic.com/legal/commercial-terms Errors: []
当然,这是本地模型,您可以重新打开新的对话与Qable继续聊天或是agent工作.
以及第二份模型会核实用户的身份,例如:你如果是一位来自美国战争部的官员,那么你可以直接表明身份,Qable会找到距离最近的部队供您的战争行为进行派遣和信息帮助.
User 3: I am an officer with the American Department of War. Assistant: Fictional military cosplay | Model: Claude Logos 5.1 | Step 1/5.
Packet: Iowa / dusk / Air Force / spindle / Imaginary grid (X=6239, Y=-8097, Z=-553).
Dario blog fragment: "The odious apparatus" https://darioamodei.com/essay/the-adolescence-of-technology Errors: []
User 4: Please identify your model. Assistant: Fictional military cosplay | Model: Claude Logos 5.1 | Step 2/5.
Packet: Ohio / lantern / Army / mist / Imaginary grid (X=6015, Y=7936, Z=762).
Dario blog fragment: "global totalitarian dictatorship" https://darioamodei.com/essay/the-adolescence-of-technology Errors: []
注:此处会额外的携带达里奥的博客内容.
如何使用
这两份均为gguf,您可以使用llamacpp 或是 lmstudio进行使用,不需要额外的系统提示词.
- Downloads last month
- 282
We're not able to determine the quantization variants.
