File size: 2,683 Bytes
7aab365
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
---
license: apache-2.0
library_name: peft
base_model: Qwen/Qwen2.5-0.5B-Instruct
tags:
  - memory
  - long-context
  - multi-session
  - consolidation
  - lora
---

# CLERK — Consolidated Ledger with Eviction and Rewrite Keys

**Fixed-budget, write-time memory consolidation for LLMs, trained (not
prompted).** CLERK maintains memory as a JSON ledger of atomic fact slots.
At every session boundary a trained policy reads (ledger, session
transcript) and emits a structured edit program — `ADD`, `UPDATE`
(supersede), `TOMBSTONE`, `EVICT` — which a deterministic reducer applies.
Reading then costs O(budget) tokens forever, regardless of history length,
and supersession is explicit rather than buried in stale text.

## Why this is new

- All published write-time consolidation is **prompt-heuristic** (Mem0,
  Zep, CUPMem); prior work shows LLM-consolidated memories corrupt over
  repeated updates. CLERK **trains** the write policy on exact supervision
  from programmatically generated gold ledger transitions.
- Learned memory policies (Memory-T1, MemAgent) act at **read time**;
  CLERK acts at write time with a **fixed budget** and learned eviction.
- Prior fixed-budget parametric memories (RMT, Infini-attention) are
  opaque token memories trained from scratch; CLERK is **interpretable,
  reducer-checked, and LoRA-scale**.

## Repository layout

```
clerk/
  common.py         slot schema, edit-op reducer, serializers, salience oracle
  prompts.py        CONSOLIDATE / ANSWER instruction formats (single source of truth)
  generator.py      seeded synthetic evolving-session generator (gold ledger states)
  build_sft.py      timelines -> SFT messages dataset (CONSOLIDATE + ANSWER mixture)
  train_sft.py      LoRA SFT via TRL (verified against trl/examples/sft_qlora)
  eval_synthetic.py held-out benchmark: clerk vs prompted-ledger/full/window/RAG
  eval_locomo.py    LoCoMo-MC10 zero-shot (Percena/locomo-mc10)
  tests_smoke.py    data-invariant checks (run before any GPU job)
run_all.sh          end-to-end pipeline
requirements.txt    pinned dependencies
paper/              the paper
```

## Reproduce

```bash
pip install -r requirements.txt
huggingface-cli login
bash run_all.sh          # ~2-3 GPU-hours on one 16GB GPU (L4/A10G class)
```

Everything is seeded (generator seed 137/991/4242; training seed 137).
Smoke tests verify: gold op programs are executable by the reducer,
budgets are never exceeded, serializers round-trip, and QA supervision
matches the memory state it was answered against.

## Results

See `paper/paper.md` and `results/` (populated by `run_all.sh`).

## License

Apache-2.0. Author: Justin Wolcott, fallnai-research.org.