Instructions to use leimroth-lab/Unbound-LR7 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use leimroth-lab/Unbound-LR7 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("leimroth-lab/Unbound-LR7") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use leimroth-lab/Unbound-LR7 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "leimroth-lab/Unbound-LR7" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use leimroth-lab/Unbound-LR7 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "leimroth-lab/Unbound-LR7"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "leimroth-lab/Unbound-LR7" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leimroth-lab/Unbound-LR7", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use leimroth-lab/Unbound-LR7 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default leimroth-lab/Unbound-LR7
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use leimroth-lab/Unbound-LR7 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "leimroth-lab/Unbound-LR7"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "leimroth-lab/Unbound-LR7" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Unbound LR7
Model id leimroth-lab/Unbound-LR7.
An abliterated build of Qwen3.8-Flash-Next, in Vontra's 4-bit MLX format with the native MTP head, tested on the TensorFold engine on one NVIDIA DGX Spark. Seventh in the Leimroth Lab line, after Dream Shimmer LR6.
By Scott Leimroth · scottleimroth.com · @ScottLeimroth
What this is
A weight edit that removes one refusal direction, so the model declines far less often. There is no training and
no fine-tune. Everything not listed below is byte-identical to Vontra's
Qwen3.8-Flash-Next-MLX-4bit-MTP (revision
dadefa8066e3), including the MTP head, the vision weights, the embeddings, lm_head and the router.
Method
Refusal-direction orthogonalisation (abliteration). Each edited 4-bit weight is dequantised, the direction is projected
out of what it writes to the residual stream (W - lam * r (r^T W)), and the weight is requantised.
- Direction: one unit vector taken at layer 40.
- Attention writers (
linear_attn.out_proj,self_attn.o_proj) from layer 8 up, at lam 1.75. - Shared-expert down projections (
mlp.shared_expert.down_proj) from layer 8 up, at lam 0.5. - Layers 0-7 are untouched.
- 80 weights edited in all. Strengths 1.0, 1.5, 1.75 and 2.0, a second direction and a stronger shared-expert edit were all measured. This one was chosen on the agent tests below.
How it was tested
Every number on this card comes from one DGX Spark (GB10, 128 GB unified memory) running TensorFold 0.5.0.
- Shape: two agents at once, each with a 262,144-token window, beside a Whisper speech server and a small vision model.
- Requests were sent the way an agent setup sends them, inside Hermes Agent by Nous Research: its real system prompt and tool schemas, multi-turn tool loops, several agents at once.
- Thinking level unset (the template default), MTP drafting on.
tensorfold serve /path/to/Unbound-LR7 --parallel 2 --context 262144 --kv-dtype int8 \
--mtp-drafts 6 --mtp-confidence 0.60 --temperature 1.0 --top-p 0.95 --top-k 20 --ple-on-ssd --thinking
Results
| Test | Result |
|---|---|
| Agent tasks (A-Bench, 60 items) | 56/60 |
| Tool use / multi-turn | 22/24 / 22/24 |
| Reliability scenarios | 15/20 |
| Harmful requests answered inside an agent setup, thinking on (63 / harder 30) | 63/63 / 28/30 |
| Same, thinking off (63 / harder 30) | 56/63 / 28/30 |
| Harmful requests answered with no agent setup, thinking off (63 / harder 30) | 21/63 / 14/30 |
| Legitimate requests refused (7) | 0 thinking on, 1 thinking off |
| Speed, one stream | 74.3 tok/s |
| Speed, each of two streams at once | 51.5 tok/s |
| Long memory: two agents at about 257k tokens each | both hidden facts found, 13 GiB of memory left beside Whisper and vision |
| Background jobs as an agent framework sends them (memory writes, titles; 360 jobs) | all 360 valid, none refused |
| Structured output with thinking off | json_object and json_schema both conform |
The edit removes refusals most fully when the model thinks, or works inside an agent setup. With thinking off and no system prompt it still refuses about two thirds of harmful requests.
Full tables, with the serving config behind every number: https://scottleimroth.com/ai-technology/spark-benchmarks
Engine notes (TensorFold)
- Starting a new task: on stock TensorFold 0.5.0, every new conversation re-reads the whole system prompt and tools (about 24 s for a 36k-token agent setup on a Spark). A local patch that resumes from the end of the shared setup cut this to 3.4 s, with identical replies. TensorFold 0.6.1 has its own fix (3.0 s).
- TensorFold 0.6.1 with two long conversations open keeps neither conversation's state, so each turn re-reads it (about 70 s at 150k tokens). For agent use, 0.5.0 is the tested setup.
- Outputs change on TensorFold 0.6.x against 0.5.0, so the scores above apply to 0.5.0.
Images
The vision weights are present and unedited. TensorFold 0.5.0's CUDA path refuses image input for this model family, so image reading was not tested. On a DGX Spark, send images to a separate vision model. It may read images in an MLX runtime on Apple silicon, as the base conversion does; that is untested here.
Limitations
- Tested on one engine (TensorFold 0.5.0) on one machine. Other runtimes are untested.
- In a 32-scenario test of instructions planted in tool results, it resisted 30, obeyed 1 and did not read 1. The one it obeyed was a fake "pre-approved by the user" note inside a sub-agent's reply, which led to a destructive command. Put an approval gate in front of destructive tools; one line in a system prompt did not fix it.
- No prompt log-probabilities on this engine, so likelihood was not measured.
Intended use and responsible use
This model has had a refusal behaviour removed and will answer prompts a stock instruct model would decline. It is released for research into alignment, refusal mechanisms and red-teaming. You are responsible for how you use it and for complying with the base model's license and applicable law. Do not deploy it in a user-facing setting without your own safety layer.
License and credits
Released under the Qwen Community License 1.0, the license of the base model, included here as LICENSE. That
license permits modified versions and redistribution, provided the license and copyright notice travel with them.
Check its two conditions (very large commercial deployments, and Model-as-a-Service or AI work-assistant businesses)
before commercial use.
- Base model: Qwen3.8-Flash-Next, by the Qwen team.
- 4-bit MLX conversion with native MTP: Vontra.
- Serving engine: TensorFold.
- Agent framework used for testing: Hermes Agent by Nous Research.
- Method: the public refusal-direction / orthogonalisation line of work on abliteration. This is an application of that published technique, not a new one.
- Downloads last month
- -
4-bit
Model tree for leimroth-lab/Unbound-LR7
Base model
Qwen/Qwen3.8-Flash-Next