FEM-ASM: Modular Language Computation

Community Article
Published October 8, 2026

An attention-free language model, reconstructive memories, executable skills—and the interfaces that connect them.

What changes when a language-model system’s memory, exact execution, and contextual computation can be constructed and updated separately?

This project explores that question through working research prototypes:

  • An attention-free Multi-Mesh language model, trained from scratch on approximately 28.6 billion prediction targets.
  • A separate roughly half-billion-parameter modular core that incorporates external contributions through typed residual assembly.
  • Reconstructive document memories that can be prepared independently under a frozen shared codec.
  • A deterministic virtual machine whose execution results enter the neural computation through explicit interfaces.
  • Controlled experiments in evidence attachment, removal, and result rendering.

The purpose is not to outperform Transformers or announce their replacement. It is to create an experimental setting in which the organization of language computation can itself be investigated.

The prototypes already make some of these questions testable. They also expose where an apparently successful component fails to produce a successful system.

Conceptual overview of FEM-ASM: fixed observations, reconstructive memory, and deterministic execution connected through a shared interface.

Conceptual overview of the research direction. Multi-Mesh explores fixed token identities, binary tiles, and glyph observations; the separate modular core investigates memory and execution interfaces. The diagram connects these ideas rather than depicting a single model’s forward pass.

Paper and artifacts

1. Why this began with input embeddings

The earlier experiments asked whether useful language modeling requires an independently trainable input vector for every token.

The research progressed from frozen visual Unicode representations to fixed minimal binary token codes, and then to larger-scale comparisons using a standard subword tokenizer.

The important distinction was between:

  • Identifying the observed token.
  • Learning how to use that token in context.

A fixed injective code preserves token identity. Shared trainable computation can learn useful representations from it without changing the token’s input coordinates.

That does not make learned embeddings useless. Nor does it locate “meaning” in a particular layer.

What it provides is a stable experimental boundary: the input coordinates remain unchanged while the rest of the system learns.

This prompted a broader question:

If token identity can be separated from its learned interpretation, what other parts of a language-model system can be given explicit interfaces and independent update procedures?

This is a research motivation, not a deduction that modularity must work. Fixed inputs do not automatically align independently constructed components, and a modular system need not use binary inputs.

The new experiments investigate those interfaces directly.

2. Beyond a single sequence of distinct layers

A conventional deep network applies a sequence of learned transformations. This is an effective design, but it need not be the only organization explored.

The Multi-Mesh prototype represents the same causal text prefix through several interacting states:

Representation Role
Fine token field Token-level computation from fixed Binary16 inputs
Structured Binary Tiles Explicit packing of short token spans
Glyph/form observations Fixed information about token appearance
Coarse history hierarchy Representations of completed spans at larger granularities
Scratch field Internal state refined during repeated computation

The fields exchange information through explicit restriction and prolongation operators. Local convolutional operators update them repeatedly.

For the identity channel, a block of 64 tokens contains:

64×16=1024=32×32 64 \times 16 = 1024 = 32 \times 32

binary coordinates.

This permits lossless packing of the input identity bits into a tile. It does not mean that the subsequent learned tile representation or document memory is lossless.

Completed-span representations are exposed only to later token positions. The prototype therefore remains an autoregressive causal language model.

What “FEM-inspired” means

The finite-element and domain-decomposition analogy suggests:

  • local components;
  • representations at different granularities;
  • explicit mappings between representations;
  • assembly of contributions into a shared state.

It does not assert that language obeys a physical differential equation or that the implementation is a finite-element discretization of one.

The Multi-Mesh model also still has parameterized stages. The later modular core is a separate architecture that reuses shared operators across iterations.

The shift is therefore not “there are no layers anymore.” It is:

The experiment is no longer restricted to a single chain of distinct transformations: interacting states, repeated operators, and external contributions become explicit design choices.

Attention could also be investigated within this broader organization. It is not the adversary of the project.

3. The attention-free model learns measurable language behavior

The released Multi-Mesh checkpoint records:

Quantity Value
Optimizer updates 109,200
Processed prediction targets 28,626,124,800
Parameter values, excluding fixed buffers 1,675,064,832
Fixed-buffer values 51,904,512

It was trained from scratch, not initialized from SmolLM2 weights.

The tokenizer comes from the SmolLM2 family. The language-model training uses the FineWeb-Edu component of SmolLM-Corpus.

A shared evaluation view

The following table brings together the reported local harness evaluations of Multi-Mesh, the earlier AB-EXT Transformer models, and two smaller SmolLM2 references.

The AB-EXT models use the same reported training-data component and tokenizer family as the Multi-Mesh experiment, but have a larger training budget and different architectures. The SmolLM2 models have their own training histories.

This is a quality comparison, not an experiment that isolates the effect of removing attention.

Metric Multi-Mesh AB-EXT Binary16 AB-EXT Learned SmolLM2-135M SmolLM2-360M
HellaSwag, normalized accuracy 30.77 52.40 57.79 43.02 56.28
ARC-Easy, accuracy 53.37 67.59 71.63 64.44 70.24
ARC-Challenge, normalized accuracy 23.55 34.04 37.63 29.61 38.05
PIQA, accuracy 62.95 70.51 72.69 68.44 71.38
WinoGrande, accuracy 49.33 55.33 58.56 52.57 59.35
MMLU, zero-shot accuracy 23.95 25.88 25.32 24.25 25.47
LAMBADA, accuracy 4.87 42.75 47.72 42.97 53.31
WikiText, word perplexity ↓ 65.02 18.58 16.50 23.14 17.12

Accuracy entries are percentages. Evaluation standard errors are omitted here for readability and are available in the model cards.

Training budgets and interpretation
Model Approximate model size Reported training budget
Multi-Mesh 1.675B parameter values 28.6B prediction targets
AB-EXT Binary16 1.711B trainable parameters Approximately 100B prediction targets
AB-EXT Learned 1.812B trainable parameters Approximately 100B prediction targets
SmolLM2-135M 135M Approximately 2T tokens
SmolLM2-360M 360M Approximately 4T tokens

The budgets are not equivalent measures of unique data or compute.

A common harness provides a useful evaluation framework, but exact comparisons also depend on task configurations, context handling, tokenization, and likelihood accounting. The released evaluation records and model cards remain the source for those details.

Different architectures, training budgets, and recipes prevent a sample-efficiency conclusion from this table.

These results concern the standalone Multi-Mesh LM, not the later modular attachment core. No external document memory or VM execution is enabled.

What is the positive result?

The Multi-Mesh network is not producing only chance-level behavior.

Its HellaSwag normalized accuracy is 30.77%, compared with a 25% uniform-choice reference, and its PIQA accuracy is 62.95%, compared with 50%. ARC-Easy and corpus likelihood provide additional evidence of learned language regularities.

Uniform-choice references are not complete benchmark baselines. Nevertheless, the results establish nontrivial behavior in an attention-free, heterogeneous multi-field architecture.

The released model also supports autoregressive text generation. Generation examples are useful for inspecting its behavior, but the benchmark results provide broader evidence than a few selected continuations.

The unevenness matters too: WinoGrande and MMLU are near chance, and LAMBADA is weak. This is a functioning language-model prototype, not a competitive general-purpose model.

The accomplishment is a working experimental substrate—not a leaderboard victory.

4. The next step: components with different life cycles

A document archive, an arithmetic procedure, and a contextual language state do not need to be created or updated in the same way.

The modular experiments give them separate roles.

Reconstructive document memory

A shared codec produces initial document-block cells. A worker then optimizes document-specific residual states while keeping the codec frozen.

The resulting cells are published as versioned artifacts with provenance and checksums. A snapshot identifies the exact versions available to a run.

This allows document-state construction to proceed separately from core training.

The stored values are still learned state: optimization has not disappeared. It has been moved into independently managed document elements.

Executable skills

A bounded deterministic VM executes typed instructions.

In the successful arithmetic protocol, the neural controller selects among five operations. Operand extraction and program construction are deterministic.

The VM then returns a result that can be checked and replayed independently of the neural output layer.

Neural coordination

Shared type-specific adapters map memory and execution observations into proposals for the core’s internal state.

There is no separate adapter for each document.

These proposals are combined through an explicit residual-assembly interface rather than requiring every external contribution to be inserted as ordinary prompt text.

That is the interface being investigated—not a claim that it is already better than text-based retrieval or tool use.

5. Why proceed with imperfect memory and a small core?

The project deliberately continued before every component was optimized.

The published memory production run contains 52,809 active element versions near a budget of 1.7 billion stored floating-point values.

A selected readout-cache audit covered 2,910 evidence blocks from 98 elements and obtained 74.97% token reconstruction accuracy.

That figure describes the audited subset, not a lossless guarantee for the entire store. Only one audited block was reconstructed completely.

The separate attachment core has approximately 505 million parameter values. Its standalone general-language scores are weak; its purpose here is to study interaction with the memory and execution interfaces.

These choices provide an initial operating point:

Can useful interaction be observed before either the memory or the coordinating model is highly capable?

They do not establish a lower bound on future performance. Larger models and better reconstruction may help, but that remains an experimental question.

This memory is not a compression result

A local cell stores 768 BF16 values for a 64-token block:

  • latent cell: 1,536 bytes;
  • 64 uint16 token IDs: 128 bytes.

That is a factor of 12 before adding coarse states, indices, caches, and metadata.

The present motivation is independently optimized and managed neural state—not compact storage relative to raw text.

A practical comparison must include raw-text retrieval, answer quality, construction cost, storage, and serving latency.

6. Attaching evidence changes behavior without changing core weights

The downloadable selected-evidence demonstration evaluates 417 examples and 2,396 target tokens with unchanged core weights.

Evidence condition Coverage Target-token NLL ↓ Token accuracy
Required evidence absent 0% 6.5716 12.10%
Required evidence available 100% 2.8073 58.72%

The observed difference is 3.7643 nats per target token.

This is a concrete positive result: changing external evidence availability changes predictive behavior without another update to the core weights.

It is not an isolated measurement of memory-scale benefit. The conditions differ in whether the required evidence is available.

The names G1p7 and G6p5 identify source configurations used to select evidence. The downloadable demonstration does not load the complete multi-billion-value stores.

Removal and restoration

A separate reported fault-control experiment detached ten required elements. Reattachment restored the baseline loss exactly under its deterministic evaluation protocol.

That is useful evidence about reversible attachment. It should not be confused with machine unlearning: removing an external element does not erase information that may already exist in the core’s parameters.

Likewise, restoration alone does not establish broad fault isolation. That requires measuring both affected and unaffected tasks.

The broader opportunity is to make component availability an explicit intervention rather than an implicit change to a large checkpoint.

7. Exact computation is not the same as exact output

One of the clearest findings concerns how a VM result enters the neural model.

A correct computed value can still be rendered incorrectly.

The experiments compare a global result vector repeated across output positions with position-specific observations.

VM return interface In-range exact Extended-range exact
Repeated global vector 8.59% 6.25%
Character-oriented positional 88.28% 35.94%
Tokenizer-aware positional 100.00% 85.94%
Token plus character positional 100.00% 88.28%

The evaluation configuration specifies 128 generation examples per range. The extended range includes, rather than strictly excludes, the training interval.

The large improvement shows that explicit correspondence between external output and neural output positions matters in this setting.

The tokenizer-aware interface supplies intended output-token identities, and the external buffer declares the output length. This is therefore an experiment in adopting and rendering an external result—not unaided arithmetic or independently learned stopping.

The reported numeric metric extracts the first signed integer. It is not strict equality of the complete generated text.

Even with these qualifications, the engineering lesson is substantive:

A reliable executor is only one part of a reliable tool-using system. The return interface can be a major source of error.

A natural next comparison is with simpler alternatives: returning plain text, directly copying verified output, or using a hybrid output path.

8. What residual assembly does—and does not do

For a proposal zez_e attached to node ii, the external residual is:

Ri=∑e:i(e)=iwe(ze−Zi)ϵ+∑e:i(e)=iwe. R_i = \frac{\sum_{e:i(e)=i} w_e(z_e-Z_i)} {\epsilon+\sum_{e:i(e)=i}w_e}.

A deterministic consistency step moves the state toward the weighted proposal mean. A bounded learned correction then updates it further.

For fixed proposals, the consistency substep has a straightforward quadratic-energy guarantee. The complete neural update does not inherit a guarantee of better language-model loss or more accurate reasoning.

Nor is vector agreement the same as factual agreement.

Two conflicting sources can produce a compromise vector without resolving which source is correct. Provenance, uncertainty, duplication, and contradiction handling remain important research questions.

The value of the explicit operator is that these interactions can be inspected and modified. The formula is not, by itself, an explanation of reasoning.

9. What the unsuccessful experiments teach us

Several experiments distinguish failures that would otherwise be hidden inside one end-to-end score:

  • Mean-pooled memory keys did not yield useful retrieval in the reported initial runs.
  • Strong candidate retrieval did not translate into successful unrestricted answer generation.
  • Correct VM execution did not ensure correct neural rendering.
  • Improved answer likelihood did not always imply exact answers.

The repaired lexical retrieval protocols achieved high recall on controlled tasks, but some used evidence-derived spans and stored offsets. Those results establish addressability under the stated protocol, not retrieval from an arbitrary user question.

In one reported QA run, greedy exact match remained zero despite successful retrieval.

These observations do not make the working components irrelevant. They identify the next interface that needs attention.

Storage, retrieval, extraction, execution, and rendering are separate experimental targets.

10. What comes next

The next useful experiments need not begin by making every component larger.

They include:

  • Question-only retrieval: remove evidence coordinates unavailable to a real user.
  • Matched raw-text baselines: test whether reconstructive states justify their cost.
  • Controlled replacement: update one memory element or executable skill while holding the core fixed.
  • Conflict handling: attach inconsistent evidence and measure how the system responds.
  • Iteration studies: vary runtime computation without assuming that more iterations improve answers.
  • Systems measurements: report active state, latency, memory footprint, and construction cost.
  • Stronger generation protocols: evaluate full-output correctness and stopping independently.

The purpose is to turn architectural freedom into measurable behavior.

Conclusion

The earlier fixed-input experiments separated token identity from independently trainable input vectors.

The present work extends the investigation to the organization of computation itself.

An attention-free Multi-Mesh prototype learns measurable language behavior. A separate modular core responds to attached evidence without changing its weights. Deterministic execution and positional return interfaces expose a further distinction between obtaining a correct result and communicating it correctly.

These are bounded results, but they are working results.

The project is not asking readers to accept a finished alternative to Transformers. It offers an experimental setting in which representations, memories, and executable components can be developed separately—and their interactions can be tested.

Independent reproductions, simpler baselines, and carefully designed counterexamples are welcome.

Acknowledgments

Thanks to Hugging Face and the SmolLM2 team for making models, tokenizer artifacts, and research resources publicly available. SmolLM2 supplies tokenizer artifacts and valuable external evaluation references; its pretrained weights were not used to initialize the models described here.

Thanks also to the contributors to SmolLM-Corpus and FineWeb-Edu for the data resources, and to EleutherAI’s Language Model Evaluation Harness contributors for the evaluation infrastructure.

A personal thank-you to Astra, Sol, Opus, Qwen, DeepSeek, and Gemini for assistance with ideas, code, log analysis, manuscript review, and encouragement through the many failed runs and debugging sessions.

AI assistance and responsibility

Generative AI tools assisted with experimental design, literature discovery, code drafting, diagnostic interpretation, and writing. The author reviewed the AI-assisted material, carried out the reported experiments and checks, and takes responsibility for the manuscript, numerical claims, and released artifacts.

The deterministic QA construction used in these experiments does not call an LLM to generate its questions or answers.

Citation

@misc{bochkov2026parametermonolithreconstructivememories,
      title={Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models}, 
      author={A. Bochkov},
      year={2026},
      eprint={2610.04012},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.04012}, 
}
Related fixed-input research
@article{bochkov2025emergent,
  title   = {Emergent Semantics Beyond Token Embeddings:
             Transformer {LM}s with Frozen Visual Unicode
             Representations},
  author  = {Bochkov, Andrey},
  journal = {Transactions on Machine Learning Research},
  year    = {2025},
  issn    = {2835-8856},
  url     = {https://openreview.net/forum?id=Odh8IynO1o}
}
@misc{bochkov2026languagemodelstrainableinput,
      title={Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes}, 
      author={A. Bochkov},
      year={2026},
      eprint={2605.09751},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2605.09751}, 
}
@misc{bochkov2026languagemodelsneedtrainable,
      title={Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale}, 
      author={A. Bochkov},
      year={2026},
      eprint={2610.04002},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.04002}, 
}

Community

Sign up or log in to comment