MMLA Stage B — Input LRARC4, static-objective arm
This is a Stage-B research weight release: three trained input hidden→hidden Rank-4 LRARC slow initializations (rho), together with the original, always-enabled semantic adapter omega. The exact MiniCPM5-1B-SFT backbone must be obtained separately. This is a finite four-program selection policy, not a standalone chat model or a completed PDSA system.
中文:本目录保存 B 阶段选定静态目标臂的三个训练后 rho 和必需的原 omega。三种子整体保留,未挑最好的一次。A/B 已完成本轮验收;权威记忆、预测写入 Q、双状态联合效应和严格 RSI 尚未由这些权重验证。
Paper
Junyi Zou and Avrova Donz. MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation, arXiv:2606.28876v4, 2026.
MMLA is a mixer-agnostic learning architecture. Its identity is not defined by the temporarily reused Transformer components. The v4 paper supplies the architectural/theoretical context. These September 28 Stage-B weights and measurements were produced after v4's September 13 evidence cutoff; they are not claimed to be experiments already reported in that version.
@misc{zou2026mmlamemorymediatedlearningarchitecture,
title={MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation},
author={Junyi Zou and Avrova Donz},
year={2026},
eprint={2606.28876},
archivePrefix={arXiv},
primaryClass={cs.CL},
note={Version 4, 14 September 2026},
url={https://arxiv.org/abs/2606.28876v4}
}
Exact components
| Component | Identity / behavior |
|---|---|
| Frozen backbone theta0 | openbmb/MiniCPM5-1B-SFT, revision a60b37f1fc409c54e1e337b0723aaac6f92dfec0, BF16 |
| Frozen omega | Original semantic LoRA, layers 16–23, q/v projections, rank 16, alpha 32; always enabled |
| Trained rho | A [4,1536] and B [1536,4], FP32; 12,288 parameters per seed |
| Input residual | z + (z @ A.T) @ B.T, before the frozen mixer, with no extra rank scaling |
| Episode Phi | Independent clone of the selected learned rho; two support-loss SGD steps at lr 0.1 |
| Episode reset | Restore that model's learned rho, including its trained B; do not zero the residual |
Omega's rank-16 frozen semantic adapter is distinct from the rank-4 input learning residual. No historical CP reader, M19 tail, output-space Phi or separate online rank-16 Phi is included.
Files in weights/:
static-2026092811.safetensorsstatic-2026092812.safetensorsstatic-2026092813.safetensorsomega.safetensors
lineage.json records original checkpoint/tensor hashes and all required base-file hashes. manifest.json records the exported file hashes. Exported tensors were reloaded on CPU and compared exactly with the original saved tensors; no backbone forward or new benchmark was run during export. These are portable weight exports; optimizer/RNG continuation state remains in the original training checkpoints.
Load and inspect
Use Python with PyTorch and safetensors. Loading the backbone additionally needs Transformers and its tokenizer dependencies. The reference run used the versions in environment.json.
from load_weights import verify_bundle, load_rho, begin_episode
verify_bundle()
rho = load_rho(seed=2026092811, device="cpu")
phi = begin_episode(rho) # independent A/B with gradients enabled
# Reset by calling begin_episode(rho) again.
The default seed is the first registered seed, not an assertion that it is the best or representative of all three. For the full frozen substrate, supply a local directory of the exact pinned upstream revision:
from load_weights import load_backbone, load_rho, begin_episode, apply_input_residual
model, tokenizer = load_backbone("/path/to/exact/MiniCPM5-1B-SFT", device="cuda:0")
phi = begin_episode(load_rho(2026092811, device="cuda:0"))
# For one episode, obtain embeddings and inject before entering the mixer:
# embeddings = model.model.embed_tokens(input_ids)
# inputs_embeds = apply_input_residual(embeddings, phi)
# model(inputs_embeds=inputs_embeds, attention_mask=mask, position_ids=positions, ...)
load_backbone verifies the original base-file and tensor identities and installs omega; the returned backbone alone does not apply rho. The code above is component loading, not a generate()/AutoModel pipeline or a reproduction of the benchmark. episodic_core.py provides the original independent fast-state and reset/update primitives. Do not wrap the frozen mixer in no_grad when differentiating Phi, reuse a post-Phi cache after an update, or share writable Phi/optimizer/RNG across episodes.
Training and measured scope
The original comparison trained static and post-adaptation outer objectives at the same input LRARC4 location: 2 arms × 3 paired seeds × 256 outer updates, batch 2 episodes. Both arms receive the same public support feedback and outer query supervision. Initial A is nonzero and initial B is zero. The adapted arm uses an identity-Jacobian first-order meta-gradient estimate; it is not an exact second-order algorithm. Only the selected static arm's three weight exports are included here.
The native SFT model scores four public DSL candidates by mean answer-plus-EOS token log probability, with softmax temperature 1. The primary endpoint is expected execution error after real-feedback adaptation, averaged within parameter groups and then across paired seeds. Development comprises 16 independent parameter groups, 64 recipe episodes, and 3 training seeds; seeds are not additional independent tasks. These are known operator/recipe families with held-out parameter groups, not unrestricted program generation.
| Objective | Post-adaptation expected error, mean ± seed SD | Real-feedback gain G | Permutation-adjusted gain D |
|---|---|---|---|
| Static — weights included | 29.30% ± 6.59 pp | 24.65 pp | 28.36 pp |
| Adapted — comparison only | 24.61% ± 15.66 pp | 36.65 pp | 35.77 pp |
G = real/reset error − real/keep error. D subtracts the corresponding independent uniform-permutation feedback gain. Both arms pass the preregistered feedback criteria. The static arm was selected by the fixed worst-seed stability rule, not by changing the primary endpoint or choosing the best seed. Adapted-objective improvement averages 4.69 pp (99% task-group bootstrap interval [1.19, 8.88] pp), but one seed reverses direction; this interval conditions on the three observed seeds.
A public support-based symbolic selector achieves 0.78125% error, better than either model arm. These results therefore demonstrate restricted feedback adaptation, not superiority over strong symbolic rules, a completed predictive writer, full PDSA, or strict RSI. The final audit remains unopened. Exact per-seed measurements and contrasts are in evaluation_summary.json.
Model tree for MMLA-ORG/MMLA-LRARC4
Base model
openbmb/MiniCPM5-1B-SFT