Instructions to use Johnny221B/memfold with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Johnny221B/memfold with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
MemFold
Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
This repository contains MemFold's selected reader LoRA adapters and matching memory components for Qwen3-4B, Qwen2.5-3B-Instruct, and Qwen2.5-7B-Instruct. The implementation and training instructions are in GitHub.
Updated Qwen3-4B checkpoint · 2026-09-29
The default PersonaMem-32K / Qwen3-4B reader is now the validation-selected more-grpo, epoch 5 (615 steps) checkpoint. Its matching compressor is unchanged. Other model bundles are unchanged.
| Split | Five-trial mean | Best trial |
|---|---|---|
| Validation | 75.2% | 86.0% |
| Test | 76.0% | 82.0% |
Both memory generation and answer decoding are greedy; trials use different fixed option permutations. 86% is a validation score, not the test score. Checkpoint selection used validation mean. Configuration and all trial scores are in OPTIMIZATION.json.
Download personamem-32k/qwen3-4b/ from the default branch for this updated reader. To reproduce the original paper release, set revision="d39842eeca8132288c5e2814fdd66733fa92fc3d" when downloading. The previous weights remain available at that fixed revision. This post-release update does not replace historical paper results.
Updated Qwen2.5-3B-Instruct checkpoint
The default personamem-32k/qwen2.5-3b/ reader is now current-ratio-epoch-4, selected by five-trial validation mean. Only the final OPD + GRPO phase was retrained; its compressor is unchanged.
| Split | Mean accuracy | Best trial |
|---|---|---|
| Previous reader validation | 60.0% | 70.0% |
| Updated reader validation | 67.6% | 76.0% |
| Updated reader test | 61.2% | 66.0% |
Memory generation and answers use greedy decoding; trials change answer-option order. Validation best scores are not test scores. The test split was evaluated after checkpoint selection. Settings and all validation results. Previous weights remain at revision 10e71a32171ccde6f3a53bc8ee5b407264321c3d. This is a post-release update, not a replacement for historical paper results.
Updated Qwen2.5-7B-Instruct checkpoint
The default personamem-32k/qwen2.5-7b/ reader is now current-ratio-epoch-5, selected by five-trial validation mean. Only the final OPD + GRPO phase was retrained; its compressor is unchanged.
| Split | Mean accuracy | Best trial |
|---|---|---|
| Previous reader validation | 72.8% | 76.0% |
| Updated reader validation | 74.4% | 78.0% |
| Updated reader test | 80.0% | 86.0% |
Memory generation and answers use greedy decoding; trials change answer-option order. Validation best scores are not test scores. The test split was evaluated after checkpoint selection. Settings and all validation results. Previous weights remain at revision c86ee86e01369478b300007d7500a1a63ea582b2. This is a post-release update, not a replacement for historical paper results.
Updated PersonaMem-128K / Qwen3-4B
Default reader: highest-lr-step-1000. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. Selected by validation mean, with test evaluated afterward.
| Split | Mean accuracy | Best trial |
|---|---|---|
| Previous reader validation | 87.03% | 89.74% |
| Updated reader validation | 89.16% | 91.94% |
| Updated reader test | 87.64% | 89.70% |
273 validation questions and 233 test questions; greedy writer and reader, five fixed option permutations. Use the matching 128K code and protocol for YaRN4.5, full inputs without truncation, and the historical answer-format normalization. Full settings and results. Previous weights remain at revision 7de39f1557b44a3cdac0dec768f31e62e0c2180f. These post-release results do not replace historical paper results.
Updated PersonaMem-128K / Qwen2.5-3B-Instruct
Default reader: current-ratio-step-600. Only the final OPD + GRPO phase was retrained; its compressor is unchanged. Selected by validation mean, with test evaluated afterward.
| Split | Mean accuracy | Best trial |
|---|---|---|
| Previous reader validation | 89.01% | 91.58% |
| Updated reader validation | 89.23% | 90.84% |
| Updated reader test | 87.64% | 89.27% |
273 validation questions and 233 test questions; greedy writer and reader, five fixed option permutations. Use the matching 128K code and protocol for YaRN4.5, full inputs without truncation, and the historical answer-format normalization. Full settings and results. Previous weights remain at revision 71a718d07b7c61b3fdf53509557e05571368ba8d. These post-release results do not replace historical paper results.
Components and names
MemFold uses one workflow:
History + query → Writer → Text memory → Compressor → Soft tokens → Reader → Answer
- Writer generates textual memory. PersonaMem uses the selected reader adapter for this step; the LoCoMo transfer configuration uses the shared Qwen3-4B writer.
- Compressor converts memory into soft tokens using a resampler and projector. PersonaMem stores this component in
bridge.pt. - Reader is the Qwen backbone with a MemFold LoRA adapter, stored in
reader/.
Bundles are named by training dataset / backbone. Display names follow MemFold · Backbone · Dataset; they do not contain research-question or run numbers. Existing filenames and configuration keys remain compatible with the published loader.
Available weights
Each dataset directory contains qwen3-4b, qwen2.5-3b, and qwen2.5-7b.
| Bundle path | Contents | Soft-memory budget |
|---|---|---|
personamem-32k/<backbone>/ |
reader/ + bridge.pt |
256 tokens |
personamem-128k/<backbone>/ |
reader/ + bridge.pt |
256 tokens |
locomo/<backbone>/ |
reader/ + compressor.pt + mapper.pt |
512 tokens per session |
The LoCoMo-trained bundles are used for LongMemEval transfer with soft memory plus bounded text. Download shared/locomo-writer-qwen3-4b/ for their shared memory writer. The budget is per session, not per full history. The session compressor and mapper must stay paired.
qwen2.5-3b and qwen2.5-7b refer to the Instruct backbones. Each bundle.json records the exact base-model ID and revision. The release includes nine reader bundles and their inference dependencies; base-model weights and benchmark inputs are obtained separately.
Quick start
Install the code
Use Python 3.11 and a CUDA-compatible PyTorch installation. This documentation targets code commit d049fd293e3232358594cba7cdae0f7a138bb58d.
git clone https://github.com/Johnny221B/memfold.git
cd memfold
git checkout d049fd293e3232358594cba7cdae0f7a138bb58d
pip install -e '.[inference]'
Download a bundle
Run from the cloned repository:
from huggingface_hub import HfApi, snapshot_download
repo_id = "Johnny221B/memfold"
revision = HfApi().model_info(repo_id).sha
snapshot_download(
repo_id=repo_id,
revision=revision,
allow_patterns=[
"personamem-32k/qwen3-4b/*",
"load_components.py", "manifest.json", "VALIDATION.json",
"LICENSE.md", "NOTICE", "licenses/*",
],
local_dir="checkpoints/memfold",
)
print("HF revision:", revision)
Download the base model at the revision specified in bundle.json, using its matching tokenizer.
Load memory components
python checkpoints/memfold/load_components.py \
--bundle checkpoints/memfold/personamem-32k/qwen3-4b --code .
This checks component loading on CPU. The helper also exposes load_reader(bundle, **model_kwargs) for PEFT loading. Answer generation needs the complete memory pipeline below.
Generate memory and evaluate
Prepare the benchmark questions and writer inputs using the data guide. Then run the common inference example:
export MODEL=/path/to/Qwen3-4B
export BUNDLE="$PWD/checkpoints/memfold/personamem-32k/qwen3-4b"
export QUESTIONS=/path/to/test/questions.jsonl
export WRITER_INPUTS=/path/to/test/writer-inputs.jsonl
export OUTPUT="$PWD/outputs/personamem-32k"
bash examples/evaluate.sh
The example calls memfold.py generate, memfold.py encode, and memfold.py evaluate. It writes individual predictions and a summary containing each trial's accuracy, the mean, and the best result. PersonaMem uses greedy decoding; the five trials change answer-option order.
PersonaMem-128K uses the same commands with its matching bundle and inputs; set EXAMPLES=233 and a supported MAX_MODEL_LEN for the longer context. PrefEval and LoCoMo/LongMemEval use the dataset-specific adapters.
Training
Use the same code entrypoint for both PersonaMem context lengths:
prepare → extract → train writer
→ train compressor → train warmup → train reasoning
→ train reader → train optimize
For example, python memfold.py train reader --help shows reader-initialization arguments. The optimization objective combines OPD and GRPO; reference KL defaults to zero. The training guide defines each procedure and its trainable components.
Verification
manifest.json records artifact hashes and the matching code revision. VALIDATION.json distinguishes the original CPU checks from the Qwen3-4B / PersonaMem-32K GPU reproduction and the subsequent code-refactor regression. These checks do not constitute a fresh benchmark of all released models and datasets.
The documentation update preserves all adapter and memory-component weight files. Historical training settings can differ from the current code defaults.
License
MemFold-authored components and helper code use MIT, subject to the upstream model terms. Qwen3-4B and Qwen2.5-7B-Instruct use Apache-2.0. Qwen2.5-3B-Instruct uses the Qwen Research License, including its non-commercial restriction. See LICENSE.md, NOTICE, and the included license texts.
- Downloads last month
- -