Raw KV-Cache Handoff Lab
This model-companion repository defines KVC1, a fail-closed container for raw
key/value cache tensor bytes, plus an explicit translation contract for moving a
contained idea into another model family's cache geometry.
KV tensors are runtime activations, not model weights. Direct reuse requires matching model/tokenizer identity, tensor geometry, RoPE semantics, layout, and sequence position. Cross-family handoff requires a trained and evaluated translator.
The first exact-model live gate now passes. A frozen
Qwen/Qwen2.5-0.5B-Instruct revision exported 48 real key/value tensors across
24 layers into KVC1. A separate process reconstructed the Transformers
DynamicCache and resumed with the same next token as uninterrupted inference;
maximum absolute fp32 logit difference was 1.537799835205078e-05 under a
declared 5e-05 tolerance. This proves same-model extraction and reinjection,
not cross-model translation or capability transfer.
The optional DOMCAP1 sidecar pack, Analytics Hotshot, contains a compact expert briefing,
trusted host-tool identifiers, a source policy, and evaluation probes. The
portfolio fetches the JSON artifact after a visitor elects to load the pack,
validates every requested tool against its local allowlist, and injects the
bounded briefing into the next request. DuckDB-Wasm—not the language model—runs
the calculation.
What this is
- A reproducible raw KV-cache serialization and compatibility prototype.
- A contract for a versioned source-to-destination translation layer.
- A small, inspectable companion artifact for a browser-hosted base model.
- A way to measure task improvement against byte and estimated token cost.
What this is not
- It is not a fine-tuned checkpoint or a copy of the base model.
- The current reference does not yet train a cross-family translator.
- It does not prove that a small model inherits a larger model's intelligence.
- It grants no credentials, arbitrary network access, SQL execution, or source-code execution.
Artifacts
reference/kvcache.py: KVC1 reference reader/writer and inspector.reference/torch_kvcache.py: tensor-exact Transformers cache adapter.reference/live_transformers_roundtrip.py: offline two-process parity gate.artifacts/qwen25-05b-parity.kvc: small real-cache parity artifact.validation/qwen25-05b-parity.receipt.json: measured parity result.schemas/: portable capability packet, translator registry, and software-run evaluation contracts.tests/: the 23-test container, schema, and tensor-adapter suite.translation/bridge-contract.schema.json: cross-family translator receipt.capability-packs/analytics-hotshot.domcap.json: optional operating sidecar.
The associated schema is capability-packs/domcap.schema.json. A host must
still maintain its own trusted tool catalog; a remote container must never be
allowed to supply executable implementations.
Base model and context accounting
The current desktop base model is
onnx-community/Qwen2.5-0.5B-Instruct.
Its configuration advertises 32,768 maximum positions. The portfolio uses a
separate conservative bound of 3,200 dynamic prompt characters for browser
reliability and reports the capability pack's estimated prompt footprint.
Research status
Version 0.2.0 is a working exact-model cache prototype. Same-model cache parity
is validated for the recorded Qwen2.5 fixture, while cross-model and comparative
base-versus-container capability scores have not yet been published. The
translator therefore remains research, not validated.
Two later experiments are specified in the portfolio research note:
- directed conflict retention for targeted catastrophic-interference tests;
- anisotropic relational retrieval, inspired by view-dependent Gaussian representations but not assumed to share their geometry or compression law.