|
Download README.md from jkillay/context-cache-lab-stage2b-extractor: direct link, hf CLI and curl.
- Browser
- Download file 5.77 kB
-
https://huggingface.co/jkillay/context-cache-lab-stage2b-extractor/resolve/main/README.md
- Command line
-
hf download hf://jkillay/context-cache-lab-stage2b-extractor/README.md
-
curl -L -o README.md https://huggingface.co/jkillay/context-cache-lab-stage2b-extractor/resolve/main/README.md
5.77 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3-4B | |
| tags: | |
| - kv-cache | |
| - context-compression | |
| - long-context | |
| - research | |
| - negative-results | |
| - reproducibility | |
| library_name: safetensors | |
| # context-cache-lab — Stage 2b memory extractor (research artifact) | |
| > ## ⚠️ This is a negative-result research artifact, not a useful model. | |
| > | |
| > It is published so that a **failed** experiment is reproducible. Its own | |
| > preregistered gates returned **INCONCLUSIVE** (Stage 2b) and **FAIL** (Stage 3). | |
| > Do not use it for anything in production. Do not read a capability claim into it. | |
| A 566M-parameter C²KV-style sidecar that compiles chunks of text into compressed, | |
| pre-RoPE KV "pages" for a **frozen** `Qwen/Qwen3-4B`. It does not generate text and | |
| cannot be used standalone — it only produces key/value state that the frozen target | |
| consumes. | |
| **Code, protocol, preregistration and all results:** | |
| https://github.com/johnathonkillaly/context-cache-lab | |
| ## What was measured | |
| | Stage | Verdict | Result | | |
| |---|---|---| | |
| | 2b — does learned compressed state beat equal-budget raw text? | 🟡 **INCONCLUSIVE** | Beat `BUDGET` at every ratio, but missed 2 of 4 frozen criteria by ~1% | | |
| | 3 — do independently compiled pages support two-page reasoning? | ❌ **FAIL** | **0/13** valid items vs NATIVE **13/13**; all 5 gates missed | | |
| **Stage 2b (single-document QA, 4096 tokens, `clean_hit`):** | |
| | ratio | this extractor | equal-budget raw text | NATIVE | | |
| |---|---|---|---| | |
| | 2× | 0.569 | 0.472 | 0.958 | | |
| | 4× | 0.389 | 0.222 | 0.958 | | |
| | 8× | 0.347 | 0.125 | 0.958 | | |
| | 16× | 0.278 | 0.069 | 0.958 | | |
| Semantic facts survived compression (0.625 at 4×) while exact strings did not — | |
| hashes collapsed to **0.000** and identifiers to 0.250, where raw text at the same | |
| budget scored 0.750. Warm reuse was 23–32× faster than native prefill with breakeven | |
| at 2 queries, but compilation costs more than a single prefill, so the first query is | |
| slower. | |
| **Stage 3** then evaluated this exact frozen checkpoint on a preregistered cross-page | |
| compositional task. It scored **zero** clean accuracy on every native-valid item, with a | |
| mean rank margin (−2.47) *worse* than no context at all (−1.10). Compiling the document | |
| jointly instead of per-page also scored zero, so this is not an independence tax. | |
| **What that failure does not establish:** Stage 3 used a different corpus (~94–100 token | |
| pages against ~256-token training chunks), so it measures insufficient transfer of *this* | |
| carrier to *that* task — not the impossibility of independently compiled memory in | |
| general. It also cannot fully separate loss of within-page fidelity from failure to | |
| combine intact relations. See the repository's Stage 3 report for the full limitations. | |
| ## Files | |
| | File | Purpose | | |
| |---|---| | |
| | `stage2b_extractor.safetensors` | **Use this.** 109 fp32 tensors, safe to load. Verified to round-trip exactly against the original. | | |
| | `stage2b_extractor.pt` | Provenance artifact only. Its SHA-256 is pinned in the repo's `results/stage3/freeze.json`, so the Stage 3 integrity audit reproduces exactly. **It is a pickle** — the repo's scripts load it with `weights_only=False`. Only use it if you need the hash to verify. | | |
| | `config.json` | `layer_share`, `n_sink`, training step, target model revision, and the original `.pt` SHA-256. | | |
| ``` | |
| original .pt sha256: b8384dd2eeaa87890a6b65e22f791183399764b93dd9a7f4e8a6e4ce13fca425 | |
| target model revision: 1cfa9a7208912126459214e8b04321603b3df60c | |
| training step: 2200 | |
| ``` | |
| ## Usage | |
| The extractor is only meaningful alongside the repository's code: | |
| ```bash | |
| git clone https://github.com/johnathonkillaly/context-cache-lab | |
| cd context-cache-lab | |
| uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]" | |
| ``` | |
| ```python | |
| from huggingface_hub import hf_hub_download | |
| from safetensors.torch import load_file | |
| from ccl.compressor import MemoryExtractor | |
| from ccl.target import TargetModel | |
| target = TargetModel("Qwen/Qwen3-4B") # frozen | |
| extractor = MemoryExtractor(target, layer_share=1, n_sink=8) | |
| extractor.load_state_dict(load_file(hf_hub_download( | |
| "jkillay/context-cache-lab-stage2b-extractor", | |
| "stage2b_extractor.safetensors"))) | |
| extractor.eval() | |
| pages = extractor.compile_chunks([target.encode(c) for c in chunks], ratio=4.0) | |
| ``` | |
| Requires `transformers>=5.0` (the legacy tuple KV-cache format is gone) and roughly | |
| 10 GB of memory for the frozen target plus the sidecar. Developed on Apple Silicon | |
| (MPS); no CUDA is assumed. | |
| ## Training | |
| 2,200 steps, ~108 minutes on an Apple M4 Max, compression ratio sampled per example | |
| (C²KV's `-Dyn` variant). The target model was frozen throughout — only the shared | |
| memory-token embedding and per-layer Q/K/V projection heads received gradients. | |
| Supervision was on answer tokens only, applied **after** page concatenation, so the | |
| extractor is pushed toward states that compose rather than states that are merely | |
| individually informative. | |
| Training data was a **synthetic** corpus generated deterministically from integer | |
| seeds, with value pools provably disjoint from the held-out evaluation draw. No | |
| scraped text and no personal data. | |
| ## License and attribution | |
| Apache-2.0. This is a derivative of [`Qwen/Qwen3-4B`](https://huggingface.co/Qwen/Qwen3-4B) | |
| (Apache-2.0) in the sense that its projections were initialized from that model's | |
| weights; **no Qwen weights are redistributed here** — only the trained sidecar. | |
| The mechanism is a reimplementation of the published design of | |
| [C²KV (Du et al., KDD 2026)](https://arxiv.org/abs/2607.17715). No C²KV source code was | |
| copied. Please cite that paper alongside this artifact. | |
| ## Citation | |
| See `CITATION.cff` in the | |
| [GitHub repository](https://github.com/johnathonkillaly/context-cache-lab). | |