Instructions to use Adapa360/HoloQwen3Growing with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Adapa360/HoloQwen3Growing with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Adapa360/HoloQwen3Growing", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Adapa360/HoloQwen3Growing", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Adapa360/HoloQwen3Growing with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Adapa360/HoloQwen3Growing" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Adapa360/HoloQwen3Growing", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Adapa360/HoloQwen3Growing
- SGLang
How to use Adapa360/HoloQwen3Growing with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Adapa360/HoloQwen3Growing" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Adapa360/HoloQwen3Growing", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Adapa360/HoloQwen3Growing" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Adapa360/HoloQwen3Growing", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Adapa360/HoloQwen3Growing with Docker Model Runner:
docker model run hf.co/Adapa360/HoloQwen3Growing
HoloQwen3 growing conversation memory
This extension retains conversations across restarts, retrieves dated evidence, and trains small private neural adapters. The original Qwen3-0.6B transformer weights remain frozen. The archive can pass 180,000 stored tokens and keep growing as disk and compute resources permit; this does not change Qwen's attention window.
Download and run
Download the complete repository to include the conversation runtime, tests and synthetic demo. Use an isolated Python environment with the pinned requirements:
hf download Adapa360/HoloQwen3Growing --local-dir HoloQwen3Growing
python -m pip install -r HoloQwen3Growing/requirements.txt
python HoloQwen3Growing/conversation.py --model HoloQwen3Growing --memory private/alice.db
Access to a private repository requires an authenticated Hugging Face account with permission to read it. Install the appropriate PyTorch 2.6.0 CPU/CUDA wheel for your system. The tested stack is Python 3.12, Transformers 4.57.6 and PyTorch 2.6.0; the scaling run used an RTX 4090 Laptop GPU.
Qualification: 190,524 generated self-play tokens, 193,440 retained tokens, 133/133 sampled recall answers, 19/19 unknown-question abstentions, and 49,152 trained private adapter parameters. See TEST_REPORT.md, validation.json and manifest.json. These are synthetic local-workload results; general production certification and broad semantic recall have not been established.
The conversation runtime enables persistence and growth. Standard Transformers
AutoModelForCausalLM loading exposes the frozen base plus the original bounded
spectral-memory API. New conversation databases start empty; the separately
labelled validation_state/synthetic-conversations.db contains the trained
synthetic demo and is not personal user memory.
Local chat
From the project folder, after the bundle has been built:
.venv/Scripts/python.exe holoqwen3_models/HoloQwen3-Growing/conversation.py --model holoqwen3_models/HoloQwen3-Growing --memory holoqwen3_models/private/alice.db
The interface retains each user message and generated response. Type
/memory-status to inspect stored tokens and neural parameter counts,
/clear-memory to remove the archive, index, adapters and active history, and
/quit to exit. Use --no-training for archive-only operation. Memory is private
to the database you choose; use a separate path for each user.
For a single query or a historical time range:
.venv/Scripts/python.exe holoqwen3_models/HoloQwen3-Growing/conversation.py --model holoqwen3_models/HoloQwen3-Growing --memory holoqwen3_models/private/alice.db --prompt "What project label did I tell you?" --until "2026-09-05T12:00:00Z"
Time arguments are inclusive and require a UTC offset. The archive records the
host's observed UTC nanoseconds and a logical timestamp max(observed, previous+1).
The logical clock persists across restarts and maintains order through wall-clock
rollback. Timestamp precision does not imply nanosecond physical clock accuracy.
Chunks from one message share an event ID and timestamp, with a separate part index.
Storage and retrieval
Messages are tokenized with the pinned Qwen tokenizer and split into bounded chunks. Token IDs are uint32 little-endian, compressed with zlib and protected by SHA-256 integrity checks. Decoding a whole event reconstructs its token stream, including UTF-8 characters that span token/chunk boundaries.
The API session.memory.export_hex(sequence) emits the compressed record as hex,
with its format, timestamp, token count and digest. Hex is a human-readable
transport shorthand: it uses two characters per byte. The disk format is binary.
Compression measurements distinguish compressed payload bytes, Fourier keys,
index overhead and total database size; hex itself provides no compression gain.
Retrieval finds lexical candidates, ranks them with a Fourier correlation of normalized hashed word features, and supplies the strongest matches to Qwen as quoted evidence. Speaker, origin and timestamps accompany every quote. Hashing and FFTs are not encryption. By Parseval's identity, this Fourier score is the spatial feature cosine; no special Fourier compression or semantic understanding is claimed. Vocabulary-overlapping paraphrases can work, but queries that use entirely different vocabulary may miss relevant memories.
The chat prompt is bounded to 4,096 input tokens plus at most 1,024 generated tokens. Retained history can be much larger because retrieval selects relevant records. This distinction is essential: archive capacity is not a 180k native transformer context window and does not prove exhaustive recall of every token.
Neural growth and training
A rank-eight residual adapter before Qwen's frozen language-model head has
2 * hidden_size * 8 = 16,384 trainable parameters. Training runs after each
10,000 newly retained tokens by default, using user-role records and replay of
early records. Assistant answers are retained as model output, not silently
promoted to user facts. The synthetic benchmark explicitly marks both roles as
synthetic; its generated user turns are not real-world truth labels.
The first adapter is reused through 180,000 stored tokens. Each subsequent block of 10,000 tokens selects another adapter slot. Accepted adapters remain stored, and retrieval routes generation to the adapter associated with the recalled record's position in the archive. Only one small adapter runs for a response. The model's neural parameter count therefore grows, while base transformer layer counts and dimensions stay fixed. Parameter growth is not proof of additional reasoning ability or of complete factual knowledge encoded in those parameters.
Each update uses 24 AdamW steps, gradient clipping and a 10% residual-norm bound. Promotion requires reduced training NLL and no more than 5% NLL deterioration on a chronological held-out slice. Both losses, target counts and parameter counts are reported. An update that fails these gates leaves the previously accepted adapter in place. The optimizer is reconstructed for each update; it is not a hidden source of retained user information. Checkable facts remain in the archive, independent of whether adapter optimization succeeds.
Private weights are safetensors bytes stored transactionally in the same SQLite
database as the conversation records. Public save_pretrained() checkpoints do
not include these weights. Use the session API/CLI to enable this extension;
ordinary AutoModelForCausalLM.generate() still performs base-model generation.
Reset and operating scope
/clear-memory deletes records and learned adapters, removes index entries,
rotates memory IDs, clears active history and the legacy spectral bank, and
vacuums SQLite with secure-delete enabled. It restores the original model's
parameter count. It cannot erase copies in external backups, filesystem snapshots,
previous exports, terminal logs, or forensic traces on storage devices. Store
private databases in access-controlled locations and protect backups separately.
The supported application is one local process and one user per session/model instance. Calls through one session are serialized. Do not share its model with another session/thread or run multiple application processes on the same private database: in-memory adapter caches are not distributed state. SQLite transactions protect archive writes, not a multi-user serving architecture. There is no claim of a service SLO, adversarial prompt-injection immunity or real-user certification.
Reproduce evaluation
.venv/Scripts/python.exe -m unittest discover -s tests -p 'test_*.py' -v
.venv/Scripts/python.exe benchmark_conversation.py --output reports/my-independent-run --tokens 190000 --seed 47321
Use a new output directory. Eight independent conversations, each with two Qwen speakers, are generated in batches. Only actual generated tokens count toward the 10k/20k/.../190k stages; external fictional facts provide an objective recall answer key and are counted separately in stored-token totals. No repeated filler is used to reach the generated-token target. Each stage checks early, middle, recent and time-filtered answers without recent chat history, base-model answers, unknown-fact abstention and retrieval of all fact stimuli. Failure stops the scale progression and preserves evidence. Restart and reset are tested at the end.
If interrupted, rerun the same command with --resume. The checkpoint stores
speaker history, exact generated-token counts and random-generator state. A
partial synthetic batch is rolled back to its last checkpoint. Failed recall
gates cannot be bypassed by resuming. The bundled report records interruptions,
source revisions and the fixes validated during this run; historical source
snapshots are retained in validation_history.
The initial 10k development run and the frozen independent scaling run use different random seeds. Source hashes in each report identify the actual tested implementation. The bundled validation report states the measured outcome and does not turn a finite synthetic benchmark into general production certification.
The validation_state folder contains the clearly labelled synthetic trained
demo database. New user sessions start empty; do not use that demo as personal
memory or distribute a private user database in its place.
To try the trained synthetic demo, copy its database to a separate, new working
path outside the model bundle, then pass that path as --memory. Ask
What access code did I give you for observatory Tamarind81370? for a fact from
the beginning of the supplied qualification run. /memory-status shows the
retained-token and learned-parameter counts. /clear-memory clears the working
copy; the bundled reference database remains available as test evidence.
- Downloads last month
- 11