Helios Holographic Qwen3-0.6B

An initially untrained online-memory extension of Qwen3-0.6B. The base model is already pretrained by Qwen; its weights are unchanged. Empty-memory outputs match the original in the tested configurations. Accepted examples can be added over time without gradient updates or automatic training on generated answers.

Status: locally validated, upload-ready bounded research candidate. The original soft-voting memory failed populated-memory qualification; the shipped exact-context policy passes both the original and a disjoint qualification suite. This remains experimental, not universally production-certified. This is not a claim of general continual-learning ability, improved reasoning, lower perplexity, faster inference, tensor-network compression or production safety across workloads.

Architecture

The standard Qwen3 decoder has 596,049,920 unique parameters, 28 layers, hidden size 1,024, 16 query heads, 8 KV heads, and 128-dimensional heads. It is reused from Transformers rather than replaced by an approximation.

The added memory holds up to 1,024 accepted context/next-token associations. Normalized final hidden states are encoded with an orthonormal real FFT. Complex amplitudes and phases retain the complete key. Zero-lag circular correlation (the Parseval-weighted complex inner product) ranks keys. The default exact mode requires rank-1 cosine similarity at least 0.999. Conflicting target tokens whose top scores differ by no more than 1e-5 cause abstention. A valid match promotes only its token to the smallest representable logit above Qwen's current maximum. No-match and empty-memory paths preserve Qwen logits exactly. memory_mix=0 disables readout; its nonzero magnitude is ignored in exact mode.

Readout happens after the final normalized hidden state and LM head, once for each requested logit position, including cached autoregressive decode steps. It performs minimal logit rank promotion on matches, not a residual-stream addition or KV-cache update. The implementation temporarily requests decoder hidden states while memory is active; long-prefill overhead has not been qualified.

This spectral key/value bank is mathematically equivalent to cosine retrieval in the original normalized space. It does not superpose unlimited memories into constant space or compress weights/KV caches. Memory is FIFO; old records are evicted at capacity. The FP32 spectral bank and token IDs occupy 4,210,688 bytes at default capacity, excluding temporary retrieval/hidden-state tensors. Each record is one accepted next-token association, not a whole fact or conversation. Duplicates consume slots and affect votes. The 1,025th write overwrites slot 0; the cursor wraps cyclically. Retrieval does not refresh age, there is no decay, and save/load preserves the next eviction position.

An independent C++ Qwen replica and experimental genuine TT/MPO layer substitution exist in the development project; neither is used by this HF checkpoint.

Load and generate

Tested: Python 3.12, PyTorch 2.6.0, Transformers 4.57.6, safetensors 0.8.0. Review the custom code before enabling trust_remote_code. On the Hub, use the published immutable revision instead of unpinned main.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

path = "./release/holographic-qwen3-0.6b"  # or your published Hub repository
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(path, fix_mistral_regex=False)
model = AutoModelForCausalLM.from_pretrained(
    path, trust_remote_code=True, dtype=dtype
).to(device).eval()
text = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Explain Fourier transforms briefly."}],
    tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
inputs = tokenizer(text, return_tensors="pt").to(device)
with torch.inference_mode():
    result = model.generate(
        **inputs, max_new_tokens=128, do_sample=True,
        temperature=0.7, top_p=0.8, top_k=20,
    )
print(tokenizer.decode(result[0, inputs.input_ids.shape[1]:], skip_special_tokens=True))

Learn only explicitly approved examples

The explicit fix_mistral_regex=False preserves the original Qwen tokenizer; Transformers 4.57.6 can otherwise emit a spurious Mistral-regex warning for locally saved non-Mistral models. Do not apply the Mistral tokenizer substitution to Qwen.

accepted = tokenizer("The approved project code is violet.", return_tensors="pt").to(device)
model.remember(**accepted, approved=True)
model.save_memory("private/alice.memory.safetensors")

# On a later session with the same checkpoint/config:
model.load_memory("private/alice.memory.safetensors")
model.reset_memory()  # clears RAM, not previously saved files/backups

For chat, tokenize the accepted complete exchange using the same template as inference. Pass a boolean target_mask of shape `[batch, sequence] to select only accepted target tokens. Padding masks are respected. Raw text and chat-formatted contexts need not retrieve the same associations; the example above is not a guarantee that a paraphrased question retrieves the fact.

generate() never changes memory. save_pretrained() never includes private memory. Explicit save_memory() writes a separate safetensors file atomically; load_memory() validates shape, token ranges, capacity and base-checkpoint identity. Do not put private memories in the model repository. Disable/clear memory before gradient training; memory keys become stale if base weights change.

Evaluation and limitations

The bundled validation.json contains exact runs, version information, timings, source hashes and limitations. manifest.json verifies release files and includes an isolated offline AutoClass load/private-memory round-trip check.

Historical baseline evidence (does not override the failures below):

  • Exact empty-memory logits and identical greedy continuations across six prompts, evaluated independently in FP32 and BF16.
  • WikiText-2-raw test prefix: first 8,192 tokens, disjoint 256-token windows, 8,160 predicted targets. Original and empty-memory model have identical perplexity: 42.52125 FP32; 42.57556 BF16. This is not full-test-set PPL and must not be compared to differently tokenized/windowed leaderboard scores.
  • Online exact-context learning: eight novel, explicitly accepted token associations each increase their target probability immediately after insertion. This is memorization evidence, not held-out generalization or broad task accuracy.
  • Independent C++ replica: all 189 teacher-forced positions across seven prompts agree in top-1 token; maximum absolute logit error 0.00018502 against FP32 Transformers. A 16-token greedy continuation matches exactly.

Final populated-memory stress: BF16, fills 8/32/128/256/512/1,024. Sampled memory decision and final top-1 accuracy are 100% at every fill. WikiText prefix perplexity remains 42.57556 with 0/8,160 triggered positions. All four exact positives trigger, 0/16 distinct/negated probes trigger, and the evicted first query abstains. The superseded soft-voting policy scored only 25% memory-vote recall at capacity, 0% final top-1, and 5/16 false triggers; that failure remains in the development evidence.

Disjoint qualification: 64/64 previously unused synthetic multi-token completions replay exactly at a fully populated 1,024-record bank (base: 0/64). The earliest evaluated record still replays, 0/72 distinct probes trigger, conflicting values for an identical context abstain, and unrelated WikiText-prefix PPL is unchanged. This proves bounded exact replay, not paraphrase, semantic recall, truth, or real-user accuracy.

The benign injection-like write test rejects False, None, "true" and 1 without mutation, but accepts the same text with approved=True. This flag is caller consent, not authentication, semantic sanitization, or a poisoning defense. No system-instruction-override success rate was measured.

Memory introduces compute overhead. Timings are measured on one RTX 4090 Laptop GPU / Intel i9-14900HX machine; there is no demonstrated speedup. TF32 is disabled. Each throughput run includes prefill and a fixed generated-token count, with warmup and three timed repetitions. See raw samples rather than treating small timing differences as statistically significant.

Several separate Windows benchmark launches intermittently hit native access violations during Python package import or one CUDA embedding call. Unchanged reruns completed with identical gates; GPU allocations cleared and no Windows display-driver reset was logged. This host/runtime instability is unresolved and is not counted as model evidence. Validate the deployment environment separately.

Unvalidated: paraphrased/semantic retrieval, long-context limits, sustained concurrent serving, multi-GPU placement, quantization, other Transformers versions, factuality/poisoning defenses and long-term forgetting/interference on real user workloads. Use separate instances for separate users; serialize updates and concurrent reads externally. No shared multi-tenant service is supplied. Acceptance coverage does not establish the original maximum-context capability for every configuration.

The model can produce incorrect or harmful text; accepted examples can also be incorrect. The memory does not guarantee truth or provide a safety boundary. Obtain consent before retaining user data and apply access controls, retention and deletion policies. Do not deploy as an unsupervised high-stakes decision system.

Provenance

Base: Qwen/Qwen3-0.6B, revision 6130ef31402718485ca4d80a6234f70d9a4cf362. Original safetensors SHA-256: f47f71177f32bcd101b7573ec9171e6a57f4f4d31148d38e382306f42996874b.

The exported BF16 checkpoint retains original parameter values and ties the embedding/output matrix, removing the redundant serialized head copy. The approximately 1.19 GB file size is not holographic or TT compression. No private usage dataset or memory state is bundled. Original weights/tokenizer are under Apache-2.0; see LICENSE and NOTICE for attribution and modification scope.

Downloads last month
88
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Adapa360/HoloQwen3

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1342)
this model