HoloQwen3 growing conversation memory

This extension retains conversations across restarts, retrieves dated evidence, and trains small private neural adapters. The original Qwen3-0.6B transformer weights remain frozen. The archive can pass 180,000 stored tokens and keep growing as disk and compute resources permit; this does not change Qwen's attention window.

Download and run

Download the complete repository to include the conversation runtime, tests and synthetic demo. Use an isolated Python environment with the pinned requirements:

hf download Adapa360/HoloQwen3Growing --local-dir HoloQwen3Growing
python -m pip install -r HoloQwen3Growing/requirements.txt
python HoloQwen3Growing/conversation.py --model HoloQwen3Growing --memory private/alice.db

Access to a private repository requires an authenticated Hugging Face account with permission to read it. Install the appropriate PyTorch 2.6.0 CPU/CUDA wheel for your system. The tested stack is Python 3.12, Transformers 4.57.6 and PyTorch 2.6.0; the scaling run used an RTX 4090 Laptop GPU.

Qualification: 190,524 generated self-play tokens, 193,440 retained tokens, 133/133 sampled recall answers, 19/19 unknown-question abstentions, and 49,152 trained private adapter parameters. See TEST_REPORT.md, validation.json and manifest.json. These are synthetic local-workload results; general production certification and broad semantic recall have not been established.

The conversation runtime enables persistence and growth. Standard Transformers AutoModelForCausalLM loading exposes the frozen base plus the original bounded spectral-memory API. New conversation databases start empty; the separately labelled validation_state/synthetic-conversations.db contains the trained synthetic demo and is not personal user memory.

Local chat

From the project folder, after the bundle has been built:

.venv/Scripts/python.exe holoqwen3_models/HoloQwen3-Growing/conversation.py --model holoqwen3_models/HoloQwen3-Growing --memory holoqwen3_models/private/alice.db

The interface retains each user message and generated response. Type /memory-status to inspect stored tokens and neural parameter counts, /clear-memory to remove the archive, index, adapters and active history, and /quit to exit. Use --no-training for archive-only operation. Memory is private to the database you choose; use a separate path for each user.

For a single query or a historical time range:

.venv/Scripts/python.exe holoqwen3_models/HoloQwen3-Growing/conversation.py --model holoqwen3_models/HoloQwen3-Growing --memory holoqwen3_models/private/alice.db --prompt "What project label did I tell you?" --until "2026-09-05T12:00:00Z"

Time arguments are inclusive and require a UTC offset. The archive records the host's observed UTC nanoseconds and a logical timestamp max(observed, previous+1). The logical clock persists across restarts and maintains order through wall-clock rollback. Timestamp precision does not imply nanosecond physical clock accuracy. Chunks from one message share an event ID and timestamp, with a separate part index.

Storage and retrieval

Messages are tokenized with the pinned Qwen tokenizer and split into bounded chunks. Token IDs are uint32 little-endian, compressed with zlib and protected by SHA-256 integrity checks. Decoding a whole event reconstructs its token stream, including UTF-8 characters that span token/chunk boundaries.

The API session.memory.export_hex(sequence) emits the compressed record as hex, with its format, timestamp, token count and digest. Hex is a human-readable transport shorthand: it uses two characters per byte. The disk format is binary. Compression measurements distinguish compressed payload bytes, Fourier keys, index overhead and total database size; hex itself provides no compression gain.

Retrieval finds lexical candidates, ranks them with a Fourier correlation of normalized hashed word features, and supplies the strongest matches to Qwen as quoted evidence. Speaker, origin and timestamps accompany every quote. Hashing and FFTs are not encryption. By Parseval's identity, this Fourier score is the spatial feature cosine; no special Fourier compression or semantic understanding is claimed. Vocabulary-overlapping paraphrases can work, but queries that use entirely different vocabulary may miss relevant memories.

The chat prompt is bounded to 4,096 input tokens plus at most 1,024 generated tokens. Retained history can be much larger because retrieval selects relevant records. This distinction is essential: archive capacity is not a 180k native transformer context window and does not prove exhaustive recall of every token.

Neural growth and training

A rank-eight residual adapter before Qwen's frozen language-model head has 2 * hidden_size * 8 = 16,384 trainable parameters. Training runs after each 10,000 newly retained tokens by default, using user-role records and replay of early records. Assistant answers are retained as model output, not silently promoted to user facts. The synthetic benchmark explicitly marks both roles as synthetic; its generated user turns are not real-world truth labels.

The first adapter is reused through 180,000 stored tokens. Each subsequent block of 10,000 tokens selects another adapter slot. Accepted adapters remain stored, and retrieval routes generation to the adapter associated with the recalled record's position in the archive. Only one small adapter runs for a response. The model's neural parameter count therefore grows, while base transformer layer counts and dimensions stay fixed. Parameter growth is not proof of additional reasoning ability or of complete factual knowledge encoded in those parameters.

Each update uses 24 AdamW steps, gradient clipping and a 10% residual-norm bound. Promotion requires reduced training NLL and no more than 5% NLL deterioration on a chronological held-out slice. Both losses, target counts and parameter counts are reported. An update that fails these gates leaves the previously accepted adapter in place. The optimizer is reconstructed for each update; it is not a hidden source of retained user information. Checkable facts remain in the archive, independent of whether adapter optimization succeeds.

Private weights are safetensors bytes stored transactionally in the same SQLite database as the conversation records. Public save_pretrained() checkpoints do not include these weights. Use the session API/CLI to enable this extension; ordinary AutoModelForCausalLM.generate() still performs base-model generation.

Reset and operating scope

/clear-memory deletes records and learned adapters, removes index entries, rotates memory IDs, clears active history and the legacy spectral bank, and vacuums SQLite with secure-delete enabled. It restores the original model's parameter count. It cannot erase copies in external backups, filesystem snapshots, previous exports, terminal logs, or forensic traces on storage devices. Store private databases in access-controlled locations and protect backups separately.

The supported application is one local process and one user per session/model instance. Calls through one session are serialized. Do not share its model with another session/thread or run multiple application processes on the same private database: in-memory adapter caches are not distributed state. SQLite transactions protect archive writes, not a multi-user serving architecture. There is no claim of a service SLO, adversarial prompt-injection immunity or real-user certification.

Reproduce evaluation

.venv/Scripts/python.exe -m unittest discover -s tests -p 'test_*.py' -v
.venv/Scripts/python.exe benchmark_conversation.py --output reports/my-independent-run --tokens 190000 --seed 47321

Use a new output directory. Eight independent conversations, each with two Qwen speakers, are generated in batches. Only actual generated tokens count toward the 10k/20k/.../190k stages; external fictional facts provide an objective recall answer key and are counted separately in stored-token totals. No repeated filler is used to reach the generated-token target. Each stage checks early, middle, recent and time-filtered answers without recent chat history, base-model answers, unknown-fact abstention and retrieval of all fact stimuli. Failure stops the scale progression and preserves evidence. Restart and reset are tested at the end.

If interrupted, rerun the same command with --resume. The checkpoint stores speaker history, exact generated-token counts and random-generator state. A partial synthetic batch is rolled back to its last checkpoint. Failed recall gates cannot be bypassed by resuming. The bundled report records interruptions, source revisions and the fixes validated during this run; historical source snapshots are retained in validation_history.

The initial 10k development run and the frozen independent scaling run use different random seeds. Source hashes in each report identify the actual tested implementation. The bundled validation report states the measured outcome and does not turn a finite synthetic benchmark into general production certification.

The validation_state folder contains the clearly labelled synthetic trained demo database. New user sessions start empty; do not use that demo as personal memory or distribute a private user database in its place.

To try the trained synthetic demo, copy its database to a separate, new working path outside the model bundle, then pass that path as --memory. Ask What access code did I give you for observatory Tamarind81370? for a fact from the beginning of the supplied qualification run. /memory-status shows the retained-token and learned-parameter counts. /clear-memory clears the working copy; the bundled reference database remains available as test evidence.

Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Adapa360/HoloQwen3Growing

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1349)
this model