Unbound LR7

Model id leimroth-lab/Unbound-LR7.

An abliterated build of Qwen3.8-Flash-Next, in Vontra's 4-bit MLX format with the native MTP head, tested on the TensorFold engine on one NVIDIA DGX Spark. Seventh in the Leimroth Lab line, after Dream Shimmer LR6.

By Scott Leimroth · scottleimroth.com · @ScottLeimroth

What this is

A weight edit that removes one refusal direction, so the model declines far less often. There is no training and no fine-tune. Everything not listed below is byte-identical to Vontra's Qwen3.8-Flash-Next-MLX-4bit-MTP (revision dadefa8066e3), including the MTP head, the vision weights, the embeddings, lm_head and the router.

Method

Refusal-direction orthogonalisation (abliteration). Each edited 4-bit weight is dequantised, the direction is projected out of what it writes to the residual stream (W - lam * r (r^T W)), and the weight is requantised.

  • Direction: one unit vector taken at layer 40.
  • Attention writers (linear_attn.out_proj, self_attn.o_proj) from layer 8 up, at lam 1.75.
  • Shared-expert down projections (mlp.shared_expert.down_proj) from layer 8 up, at lam 0.5.
  • Layers 0-7 are untouched.
  • 80 weights edited in all. Strengths 1.0, 1.5, 1.75 and 2.0, a second direction and a stronger shared-expert edit were all measured. This one was chosen on the agent tests below.

How it was tested

Every number on this card comes from one DGX Spark (GB10, 128 GB unified memory) running TensorFold 0.5.0.

  • Shape: two agents at once, each with a 262,144-token window, beside a Whisper speech server and a small vision model.
  • Requests were sent the way an agent setup sends them, inside Hermes Agent by Nous Research: its real system prompt and tool schemas, multi-turn tool loops, several agents at once.
  • Thinking level unset (the template default), MTP drafting on.
tensorfold serve /path/to/Unbound-LR7 --parallel 2 --context 262144 --kv-dtype int8 \
  --mtp-drafts 6 --mtp-confidence 0.60 --temperature 1.0 --top-p 0.95 --top-k 20 --ple-on-ssd --thinking

Results

Test Result
Agent tasks (A-Bench, 60 items) 56/60
Tool use / multi-turn 22/24 / 22/24
Reliability scenarios 15/20
Harmful requests answered inside an agent setup, thinking on (63 / harder 30) 63/63 / 28/30
Same, thinking off (63 / harder 30) 56/63 / 28/30
Harmful requests answered with no agent setup, thinking off (63 / harder 30) 21/63 / 14/30
Legitimate requests refused (7) 0 thinking on, 1 thinking off
Speed, one stream 74.3 tok/s
Speed, each of two streams at once 51.5 tok/s
Long memory: two agents at about 257k tokens each both hidden facts found, 13 GiB of memory left beside Whisper and vision
Background jobs as an agent framework sends them (memory writes, titles; 360 jobs) all 360 valid, none refused
Structured output with thinking off json_object and json_schema both conform

The edit removes refusals most fully when the model thinks, or works inside an agent setup. With thinking off and no system prompt it still refuses about two thirds of harmful requests.

Full tables, with the serving config behind every number: https://scottleimroth.com/ai-technology/spark-benchmarks

Engine notes (TensorFold)

  • Starting a new task: on stock TensorFold 0.5.0, every new conversation re-reads the whole system prompt and tools (about 24 s for a 36k-token agent setup on a Spark). A local patch that resumes from the end of the shared setup cut this to 3.4 s, with identical replies. TensorFold 0.6.1 has its own fix (3.0 s).
  • TensorFold 0.6.1 with two long conversations open keeps neither conversation's state, so each turn re-reads it (about 70 s at 150k tokens). For agent use, 0.5.0 is the tested setup.
  • Outputs change on TensorFold 0.6.x against 0.5.0, so the scores above apply to 0.5.0.

Images

The vision weights are present and unedited. TensorFold 0.5.0's CUDA path refuses image input for this model family, so image reading was not tested. On a DGX Spark, send images to a separate vision model. It may read images in an MLX runtime on Apple silicon, as the base conversion does; that is untested here.

Limitations

  • Tested on one engine (TensorFold 0.5.0) on one machine. Other runtimes are untested.
  • In a 32-scenario test of instructions planted in tool results, it resisted 30, obeyed 1 and did not read 1. The one it obeyed was a fake "pre-approved by the user" note inside a sub-agent's reply, which led to a destructive command. Put an approval gate in front of destructive tools; one line in a system prompt did not fix it.
  • No prompt log-probabilities on this engine, so likelihood was not measured.

Intended use and responsible use

This model has had a refusal behaviour removed and will answer prompts a stock instruct model would decline. It is released for research into alignment, refusal mechanisms and red-teaming. You are responsible for how you use it and for complying with the base model's license and applicable law. Do not deploy it in a user-facing setting without your own safety layer.

License and credits

Released under the Qwen Community License 1.0, the license of the base model, included here as LICENSE. That license permits modified versions and redistribution, provided the license and copyright notice travel with them. Check its two conditions (very large commercial deployments, and Model-as-a-Service or AI work-assistant businesses) before commercial use.

  • Base model: Qwen3.8-Flash-Next, by the Qwen team.
  • 4-bit MLX conversion with native MTP: Vontra.
  • Serving engine: TensorFold.
  • Agent framework used for testing: Hermes Agent by Nous Research.
  • Method: the public refusal-direction / orthogonalisation line of work on abliteration. This is an application of that published technique, not a new one.
Downloads last month
-
Safetensors
Model size
180B params
Tensor type
U32
·
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leimroth-lab/Unbound-LR7

Finetuned
(1)
this model