MMLA Stage B — Input LRARC4, static-objective arm

This is a Stage-B research weight release: three trained input hidden→hidden Rank-4 LRARC slow initializations (rho), together with the original, always-enabled semantic adapter omega. The exact MiniCPM5-1B-SFT backbone must be obtained separately. This is a finite four-program selection policy, not a standalone chat model or a completed PDSA system.

中文:本目录保存 B 阶段选定静态目标臂的三个训练后 rho 和必需的原 omega。三种子整体保留,未挑最好的一次。A/B 已完成本轮验收;权威记忆、预测写入 Q、双状态联合效应和严格 RSI 尚未由这些权重验证。

Paper

Junyi Zou and Avrova Donz. MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation, arXiv:2606.28876v4, 2026.

MMLA is a mixer-agnostic learning architecture. Its identity is not defined by the temporarily reused Transformer components. The v4 paper supplies the architectural/theoretical context. These September 28 Stage-B weights and measurements were produced after v4's September 13 evidence cutoff; they are not claimed to be experiments already reported in that version.

@misc{zou2026mmlamemorymediatedlearningarchitecture,
  title={MMLA: Memory-Mediated Learning Architecture for Predictive Dual-State Adaptation},
  author={Junyi Zou and Avrova Donz},
  year={2026},
  eprint={2606.28876},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  note={Version 4, 14 September 2026},
  url={https://arxiv.org/abs/2606.28876v4}
}

Exact components

Component Identity / behavior
Frozen backbone theta0 openbmb/MiniCPM5-1B-SFT, revision a60b37f1fc409c54e1e337b0723aaac6f92dfec0, BF16
Frozen omega Original semantic LoRA, layers 16–23, q/v projections, rank 16, alpha 32; always enabled
Trained rho A [4,1536] and B [1536,4], FP32; 12,288 parameters per seed
Input residual z + (z @ A.T) @ B.T, before the frozen mixer, with no extra rank scaling
Episode Phi Independent clone of the selected learned rho; two support-loss SGD steps at lr 0.1
Episode reset Restore that model's learned rho, including its trained B; do not zero the residual

Omega's rank-16 frozen semantic adapter is distinct from the rank-4 input learning residual. No historical CP reader, M19 tail, output-space Phi or separate online rank-16 Phi is included.

Files in weights/:

  • static-2026092811.safetensors
  • static-2026092812.safetensors
  • static-2026092813.safetensors
  • omega.safetensors

lineage.json records original checkpoint/tensor hashes and all required base-file hashes. manifest.json records the exported file hashes. Exported tensors were reloaded on CPU and compared exactly with the original saved tensors; no backbone forward or new benchmark was run during export. These are portable weight exports; optimizer/RNG continuation state remains in the original training checkpoints.

Load and inspect

Use Python with PyTorch and safetensors. Loading the backbone additionally needs Transformers and its tokenizer dependencies. The reference run used the versions in environment.json.

from load_weights import verify_bundle, load_rho, begin_episode

verify_bundle()
rho = load_rho(seed=2026092811, device="cpu")
phi = begin_episode(rho)  # independent A/B with gradients enabled
# Reset by calling begin_episode(rho) again.

The default seed is the first registered seed, not an assertion that it is the best or representative of all three. For the full frozen substrate, supply a local directory of the exact pinned upstream revision:

from load_weights import load_backbone, load_rho, begin_episode, apply_input_residual

model, tokenizer = load_backbone("/path/to/exact/MiniCPM5-1B-SFT", device="cuda:0")
phi = begin_episode(load_rho(2026092811, device="cuda:0"))
# For one episode, obtain embeddings and inject before entering the mixer:
# embeddings = model.model.embed_tokens(input_ids)
# inputs_embeds = apply_input_residual(embeddings, phi)
# model(inputs_embeds=inputs_embeds, attention_mask=mask, position_ids=positions, ...)

load_backbone verifies the original base-file and tensor identities and installs omega; the returned backbone alone does not apply rho. The code above is component loading, not a generate()/AutoModel pipeline or a reproduction of the benchmark. episodic_core.py provides the original independent fast-state and reset/update primitives. Do not wrap the frozen mixer in no_grad when differentiating Phi, reuse a post-Phi cache after an update, or share writable Phi/optimizer/RNG across episodes.

Training and measured scope

The original comparison trained static and post-adaptation outer objectives at the same input LRARC4 location: 2 arms × 3 paired seeds × 256 outer updates, batch 2 episodes. Both arms receive the same public support feedback and outer query supervision. Initial A is nonzero and initial B is zero. The adapted arm uses an identity-Jacobian first-order meta-gradient estimate; it is not an exact second-order algorithm. Only the selected static arm's three weight exports are included here.

The native SFT model scores four public DSL candidates by mean answer-plus-EOS token log probability, with softmax temperature 1. The primary endpoint is expected execution error after real-feedback adaptation, averaged within parameter groups and then across paired seeds. Development comprises 16 independent parameter groups, 64 recipe episodes, and 3 training seeds; seeds are not additional independent tasks. These are known operator/recipe families with held-out parameter groups, not unrestricted program generation.

Objective Post-adaptation expected error, mean ± seed SD Real-feedback gain G Permutation-adjusted gain D
Static — weights included 29.30% ± 6.59 pp 24.65 pp 28.36 pp
Adapted — comparison only 24.61% ± 15.66 pp 36.65 pp 35.77 pp

G = real/reset error − real/keep error. D subtracts the corresponding independent uniform-permutation feedback gain. Both arms pass the preregistered feedback criteria. The static arm was selected by the fixed worst-seed stability rule, not by changing the primary endpoint or choosing the best seed. Adapted-objective improvement averages 4.69 pp (99% task-group bootstrap interval [1.19, 8.88] pp), but one seed reverses direction; this interval conditions on the three observed seeds.

A public support-based symbolic selector achieves 0.78125% error, better than either model arm. These results therefore demonstrate restricted feedback adaptation, not superiority over strong symbolic rules, a completed predictive writer, full PDSA, or strict RSI. The final audit remains unopened. Exact per-seed measurements and contrasts are in evaluation_summary.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MMLA-ORG/MMLA-LRARC4

Adapter
(4)
this model

Paper for MMLA-ORG/MMLA-LRARC4