Instructions to use chrisuthe/ha-local-helper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use chrisuthe/ha-local-helper with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chrisuthe/ha-local-helper with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chrisuthe/ha-local-helper:Q8_0 # Run inference directly in the terminal: llama cli -hf chrisuthe/ha-local-helper:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chrisuthe/ha-local-helper:Q8_0 # Run inference directly in the terminal: llama cli -hf chrisuthe/ha-local-helper:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chrisuthe/ha-local-helper:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf chrisuthe/ha-local-helper:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chrisuthe/ha-local-helper:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf chrisuthe/ha-local-helper:Q8_0
Use Docker
docker model run hf.co/chrisuthe/ha-local-helper:Q8_0
- LM Studio
- Jan
- Ollama
How to use chrisuthe/ha-local-helper with Ollama:
ollama run hf.co/chrisuthe/ha-local-helper:Q8_0
- Unsloth Desktop
- Pi
How to use chrisuthe/ha-local-helper with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chrisuthe/ha-local-helper:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "chrisuthe/ha-local-helper:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use chrisuthe/ha-local-helper with Docker Model Runner:
docker model run hf.co/chrisuthe/ha-local-helper:Q8_0
- Lemonade
How to use chrisuthe/ha-local-helper with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chrisuthe/ha-local-helper:Q8_0
Run and chat with the model
lemonade run user.ha-local-helper-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use chrisuthe/ha-local-helper with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chrisuthe/ha-local-helper:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default chrisuthe/ha-local-helper:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use chrisuthe/ha-local-helper with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chrisuthe/ha-local-helper:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "chrisuthe/ha-local-helper:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
HA Local Helper (Qwen3-0.6B)
A small model that sits in front of Home Assistant's voice pipeline and emits validated device-control tool calls, reports entity state, or declines.
It handles the local half of a two-tier assistant: copying entity names verbatim, actuating devices, and reading state. Anything needing world knowledge or web search is out of scope and belongs to a larger model.
Status: scaffolding. No weights published yet. Weights land when the current training run completes and passes the real-registry gate described below. The findings here are already measured and are the reason the run is configured the way it is.
Target contract
Home Assistant 2026.9.3, which differs from earlier releases in ways that break a model trained against them:
- per-integration
LLMToolsplatforms merged into one API, so tool names arrive namespaced:intent__HassTurnOn,light__HassLightSet - a
Static Context:header listing entities - no entity state in the prompt โ state must be fetched with
homeassistant__GetLiveContext homeassistant__GetLiveContextfilters byname,domainandarea, and its description says to combine them
The prompt shape is randomised across four axes during training (header, whether states are present, whether tool names are namespaced, whether GetLiveContext is offered) because an earlier model trained on exactly one shape broke when Home Assistant changed it.
Two findings worth carrying
LoRA, not a full fine-tune
Three models were trained on near-identical corpora:
| copies entity names verbatim | reports state | recipe | |
|---|---|---|---|
| run 1 | yes | no | LoRA |
| run 2 | no | yes | full fine-tune, loss 0.022 |
| run 3 | no | yes | full fine-tune, loss 0.022 |
The full fine-tunes substituted training vocabulary for unseen names: asked to turn off a light whose name contained a word occurring zero times in training, they emitted the name of a different entity whose words occurred tens of thousands of times.
Verbatim copying is a pretrained capability. LoRA's constraint on how far the weights can move preserves it; a full fine-tune driven to very low loss overwrites it. "Full fine-tune on the bigger GPU" is a downgrade for this task, not an upgrade.
Training loss does not separate good from bad here. Run 1 reached 0.01-0.02 and copies correctly; runs 2 and 3 reached 0.022 and do not. Checkpoints are selected on the behavioural gate, never on loss.
The <think></think> scaffold is a deployment requirement
Training renders every row with enable_thinking=False, which the Qwen3 chat
template turns into <think>\n\n</think>\n\n at the head of the assistant
turn. The model learns that its answer begins immediately after that scaffold.
Servers that do not pass enable_thinking omit it, and the model then begins
generating four tokens earlier than it ever did in training. Four tokens is
enough to break it outright: on one measurement, five of six real-registry
utterances named the wrong entity, and removing the same four tokens by hand
reproduced it exactly โ the model emits the correct call followed by a
spurious second one, and the serving path surfaces the wrong one.
So any GGUF built from this model must carry the scaffold unconditionally
in its embedded chat template. Passing think: false per request also works,
but clients cannot be relied on to send it.
Evaluation
Scored against registries exported from real Home Assistant instances, on
whichever server actually serves the model โ not against a holdout built by
the same generator as the training data, and not through transformers.
Both halves of that matter, and both were learned the hard way:
- A generated holdout reported 92.9% and 94.2% name accuracy for the two models that named invented entities in production. A check that shares the training data's assumptions can only confirm them.
transformersin float32 reported state reporting and compound commands working. Neither works on any GGUF build, on either server. Quantization, merge dtype, output dtype, samplers and prompt shape were each measured and each ruled out; every served build agrees with every other and onlytransformersdiffers.
A behaviour that only appears on the training-time runtime is not a behaviour the system has.
The gate's first check, applied to every emitted call regardless of what the
case expected: no emitted entity name or area may be absent from the
registry. Name comparison follows Home Assistant exactly โ
name.strip().casefold(), so case and surrounding whitespace are forgiven and
punctuation is not.
Known limitation in the published corpus design
Entity names in the training corpus are devices in rooms. Place, transit and world words appeared in the corpus only in rows where the correct answer is to decline, and never as part of an entity name โ eleven occurrences, eleven declines, zero counterexamples.
The effect is that an entity named after a place is declined rather than looked up, however plainly it sits in the registry: a transit-line sensor was refused while a temperature sensor in the same registry was fetched happily, and removing the place word restored it.
Real registries are full of such names โ transit lines, weather stations, flight trackers, bin collections. The corpus now includes world vocabulary in entity names to supply the missing counterexample. Anyone building a similar corpus should assume entity names will be classified by vocabulary unless taught otherwise.
Attribution
Training data derives from:
- ha-voice-test-suite โ MIT
- Home Assistant
intentsโ CC BY 4.0, attribution required
Base model: Qwen/Qwen3-0.6B.
- Downloads last month
- 165
8-bit