HA Local Helper (Qwen3-0.6B)

A small model that sits in front of Home Assistant's voice pipeline and emits validated device-control tool calls, reports entity state, or declines.

It handles the local half of a two-tier assistant: copying entity names verbatim, actuating devices, and reading state. Anything needing world knowledge or web search is out of scope and belongs to a larger model.

Status: scaffolding. No weights published yet. Weights land when the current training run completes and passes the real-registry gate described below. The findings here are already measured and are the reason the run is configured the way it is.

Target contract

Home Assistant 2026.9.3, which differs from earlier releases in ways that break a model trained against them:

  • per-integration LLMTools platforms merged into one API, so tool names arrive namespaced: intent__HassTurnOn, light__HassLightSet
  • a Static Context: header listing entities
  • no entity state in the prompt โ€” state must be fetched with homeassistant__GetLiveContext
  • homeassistant__GetLiveContext filters by name, domain and area, and its description says to combine them

The prompt shape is randomised across four axes during training (header, whether states are present, whether tool names are namespaced, whether GetLiveContext is offered) because an earlier model trained on exactly one shape broke when Home Assistant changed it.

Two findings worth carrying

LoRA, not a full fine-tune

Three models were trained on near-identical corpora:

copies entity names verbatim reports state recipe
run 1 yes no LoRA
run 2 no yes full fine-tune, loss 0.022
run 3 no yes full fine-tune, loss 0.022

The full fine-tunes substituted training vocabulary for unseen names: asked to turn off a light whose name contained a word occurring zero times in training, they emitted the name of a different entity whose words occurred tens of thousands of times.

Verbatim copying is a pretrained capability. LoRA's constraint on how far the weights can move preserves it; a full fine-tune driven to very low loss overwrites it. "Full fine-tune on the bigger GPU" is a downgrade for this task, not an upgrade.

Training loss does not separate good from bad here. Run 1 reached 0.01-0.02 and copies correctly; runs 2 and 3 reached 0.022 and do not. Checkpoints are selected on the behavioural gate, never on loss.

The <think></think> scaffold is a deployment requirement

Training renders every row with enable_thinking=False, which the Qwen3 chat template turns into <think>\n\n</think>\n\n at the head of the assistant turn. The model learns that its answer begins immediately after that scaffold.

Servers that do not pass enable_thinking omit it, and the model then begins generating four tokens earlier than it ever did in training. Four tokens is enough to break it outright: on one measurement, five of six real-registry utterances named the wrong entity, and removing the same four tokens by hand reproduced it exactly โ€” the model emits the correct call followed by a spurious second one, and the serving path surfaces the wrong one.

So any GGUF built from this model must carry the scaffold unconditionally in its embedded chat template. Passing think: false per request also works, but clients cannot be relied on to send it.

Evaluation

Scored against registries exported from real Home Assistant instances, on whichever server actually serves the model โ€” not against a holdout built by the same generator as the training data, and not through transformers.

Both halves of that matter, and both were learned the hard way:

  • A generated holdout reported 92.9% and 94.2% name accuracy for the two models that named invented entities in production. A check that shares the training data's assumptions can only confirm them.
  • transformers in float32 reported state reporting and compound commands working. Neither works on any GGUF build, on either server. Quantization, merge dtype, output dtype, samplers and prompt shape were each measured and each ruled out; every served build agrees with every other and only transformers differs.

A behaviour that only appears on the training-time runtime is not a behaviour the system has.

The gate's first check, applied to every emitted call regardless of what the case expected: no emitted entity name or area may be absent from the registry. Name comparison follows Home Assistant exactly โ€” name.strip().casefold(), so case and surrounding whitespace are forgiven and punctuation is not.

Known limitation in the published corpus design

Entity names in the training corpus are devices in rooms. Place, transit and world words appeared in the corpus only in rows where the correct answer is to decline, and never as part of an entity name โ€” eleven occurrences, eleven declines, zero counterexamples.

The effect is that an entity named after a place is declined rather than looked up, however plainly it sits in the registry: a transit-line sensor was refused while a temperature sensor in the same registry was fetched happily, and removing the place word restored it.

Real registries are full of such names โ€” transit lines, weather stations, flight trackers, bin collections. The corpus now includes world vocabulary in entity names to supply the missing counterexample. Anyone building a similar corpus should assume entity names will be classified by vocabulary unless taught otherwise.

Attribution

Training data derives from:

  • ha-voice-test-suite โ€” MIT
  • Home Assistant intents โ€” CC BY 4.0, attribution required

Base model: Qwen/Qwen3-0.6B.

Downloads last month
165
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for chrisuthe/ha-local-helper

Finetuned
Qwen/Qwen3-0.6B
Adapter
(644)
this model