Instructions to use Elda-AI/memory-resoner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Elda-AI/memory-resoner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="Elda-AI/memory-resoner", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("feature-extraction", model="Elda-AI/memory-resoner", trust_remote_code=True)# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True, device_map="auto")Access to memory-resoner
Released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.
By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.
Log in or Sign Up to review the conditions and access this model content.
memory-resoner β the reference step of conversational memory
A conversation can only be stored if you know what γκ·Έ λλ€γ meant. memory-resoner reads the turns of a conversation and says, for each mention, which earlier mention it refers to β so the layer above can write a memory entry that points at the right words.
"ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ . κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?"
mentions ν©μ λμ (s0, w0β0)
κ·Έ λλ€λ (s1, w0β1)
chains [ν©μ λμ Β· κ·Έ λλ€λ]
version 0.2.6 Β· 308M parameters Β· fp32 Β· 33 ms per sentence on CPU
when the mentions are supplied (124 ms when it finds them itself).
Where it sits
memory-resoner is one stage of a conversational stack, and it is multilingual β trained on Korean and English, measured on Japanese, and it transfers to languages it never saw. Each stage answers one question and hands the next stage a span, never prose.
utterance
β
ββ Elda-AI/intenter what was said, and what kind of thing it is ββ spans come from here
ββ slot extraction which of those the user said about themselves
β
βΌ
β
memory-resoner β
which earlier mention does this one refer to
β
βΌ
memory write a pointer into the user's own words β checkable, not generated
Where the spans come from. In the Elda stack, mention spans are produced upstream by Elda-AI/intenter and this model consumes them; its card is the reference for how a span is drawn, which types exist, and what the channels mean. Used standalone, memory-resoner will find its own mentions instead.
β Those two paths are not equally good, and the gap is large. On KLUE-WoS dialogue reference, this version scores 0.5640 when the mentions are supplied and 0.2319 when it has to find them itself. Detection and linking are two abilities and the end-to-end number is their product β so if you already have spans upstream, pass them in.
β The two sides count spans in different units, and the conversion is yours to make. Upstream spans are character offsets with the particle left outside (γν©μ λγ). This model works in μ΄μ (word) indices, and a μ΄μ contains its particle, so the same mention is γν©μ λμγ here. Neither is wrong; they are different units. Align on the μ΄μ that contains the upstream span.
What it returns. Span indices and chains. Not a rewritten sentence, not a summary β an index into the words you sent, which the layer above can verify before storing anything.
What it does not decide. Whether a fact is true, whether it is worth keeping, how long it lives. Those belong to the system around it. This stage answers one question and stops.
Why a small encoder
| this model | |
|---|---|
| Parameters | 308M |
| Latency per sentence, CPU β linking, mentions given | p50 33 ms |
| Latency per sentence, CPU β finding mentions and linking | p50 124 ms |
| Output | span index β checkable byte-for-byte |
| Determinism | same input, same chains |
Memory is written on every turn. A stage that runs that often has to be cheap, and its answer has to be something the next stage can check rather than trust. A span index is both.
Output contract
input sentences, in order β a string per sentence, or a list of μ΄μ
output mentions Β· antecedents Β· chains (links closed transitively)
spans inclusive word (μ΄μ ) indices β Korean particles stay attached (see the note above)
window 6 previous sentences of context Β· 40 previous mentions as candidates
A mention with no antecedent opens a chain of its own β that is how the next stage learns it has seen a new entity, so singleton chains are kept rather than dropped.
Usage
from transformers import AutoModel, AutoTokenizer
m = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval()
tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True)
out = m.coref(["ν©μ λμ κ°κ²λ₯Ό μ΄μ΄μ .",
"κ·Έ λλ€λ μ λμΈκ΅¬κ° λ§λμ ?"], tok)
out["mentions"] # [{'sent': 0, 'words': [0, 0], 'text': 'ν©μ λμ'}, ...]
out["antecedents"] # [None, 0] β per mention: what it points back at
out["chains"] # [[0, 1]]
| call | does |
|---|---|
m.coref(sentences, tok) |
finds the mentions and links them |
m.mentions(sentences, tok) |
finds mentions only β a span provider for another linker |
m.resolve(units, mentions, tok) |
links only β you supply the mentions |
m.coref_last(sentences, tok) |
β serving shape β only the newest turn's mentions are returned, with the earlier turns kept as context |
m.resolve_last(sentences, mentions, tok) |
β serving shape, mentions supplied β the same, but you hand in the spans |
The two _last calls are what a live conversation wants: one turn arrives, you want answers for
that turn, and the turns before it are context rather than output. They take the whole
conversation and cut the window themselves, so the caller never has to know how long the window is.
β On
coref_last, do not count withantecedentsThe two calls differ here, and it matters.
resolve_lastreturns all the mentions you gave it, so itsantecedents[i]indexes that same list and reads normally.coref_lastreturns only the newest turn's mentions β so when the antecedent sits in an earlier turn, which on a live conversation is almost always,antecedents[i]isNone. There,Nonewears one face for two different facts: "it pointed at something outside this list" and "it pointed at nothing", and a consumer that counts non-Nonegets a structural zero even when the model is working.
field what it is use it for linked[i]β bool β did this mention resolve at all counting, on either call n_linkedsum of the above counting antecedent_mentions[i]the antecedent itself ( {sent, words, text}), orNonereading the answer antecedents[i]index into the returned list, or Nonesafe on resolve_last; oncoref_lastonly for within-turn links
resolve_lastalso reports what it had to drop or fix, and none of it is thrown away silently:mentions_outside_windowΒ·mentions_unencodableΒ·window_truncatedΒ·mentions_out_of_order. β Whenmentions_out_of_order > 0the layer re-sorted your mentions into reading order, and thenantecedents[i] > ican occur β the index is into the list you gave, not into reading order. If that count is zero, "an antecedent index is always smaller than its own" holds.
Notes that matter in practice:
- fp32. The config pins it. In bf16 near-ties flip and chains change.
- Pass sentences, not a paragraph. The sentence boundary is what the window is counted in.
- Send the resolved span downstream, not the reference. The output is an index into the words you sent, so whatever consumes it can verify what it was handed.
- Strip the particle at display time, not at span time. Keeping it inside the μ΄μ is what makes the span a plain index into your own input; trimming belongs to whatever renders the value.
Measurements
Public multilingual coreference, held out from training; document overlap with the training split is 0 and the benchmark script asserts it. Both scripts ship in this repository.
β Korean is one of the languages here, not the only one. If your pipeline is Korean-only, measure before you swap: reaching other languages costs some Korean headroom, and the trade is visible below. The two scripts here let you check that against your own data rather than take our word for it.
pip install torch transformers datasets
python bench_corefud_ko.py --model Elda-AI/memory-resoner
Mention detection β boundaries must match exactly
| precision | recall | F1 |
|---|---|---|
| 0.7917 | 0.8087 | 0.8001 |
Linking, mentions given
| scored mentions | 10263 |
| accuracy | 0.8803 |
| non-NULL accuracy | 0.7045 (n=2721) |
| deictic, non-NULL | 0.7159 (n=271) |
| reference β always NULL | 0.7349 |
| reference β always most-recent | 0.0448 |
Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores 0.7349; the rows that require actually linking are the ones in bold.
End-to-end β (mention, antecedent) pairs, the model finding its own mentions
| precision | recall | F1 |
|---|---|---|
| 0.6188 | 0.6661 | 0.6415 |
Every run first pushes the gold answers back through the scorer and asserts they survive intact (1.0000). A scorer that cannot return its own gold is not grading anything.
On dialogue
The numbers above are prose. On dialogue β references mined from a public multi-turn Korean corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen by us β this version resolves γκ·Έ μλΉγ-style references at 0.5640 when the mentions are supplied, and 0.2319 when it has to find them itself.
python bench_wos_references.py --model Elda-AI/memory-resoner
β That difference is the whole point of the two paths. End-to-end is detection times linking, so a pipeline that already has spans should pass them in rather than ask for them again. β These are measured with candidates mined by the gold rule, which makes them a perfect-upstream ceiling β a real intenter will score lower. These numbers are for conversation; on plain prose the same model reads reference less reliably.
What comes next
Conversational memory here is built in three layers. They are not future versions of this model β they are different kinds of component, and saying so is the point.
| what it is for | how it is built | |
|---|---|---|
| Layer 1 Β· conversation memory | what was said in this conversation, and what refers to what | β this model + deterministic assembly |
| Layer 2 | blackbox β how a memory is addressed and kept apart from every other conversation | not described here |
| Layer 3 Β· authority and persona | whether something holds and on what grounds, and what an agent is like across conversations | a deterministic VM, and a decoder used off the request path |
Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where memories are separated from one another, and it is the part of the system we keep closed. Layer 3 is where facts acquire grounds and where a persona is formed, and it is built from two very different machines β one that must be exact, and one that must be fluent.
The VM β the exact one. Memory is written in a small closed language and executed by a deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and the same inputs always produce the same store. Reads are queries against that store, not recollection. This is what makes memory auditable rather than merely plausible, and it is why the per-turn path carries no generation at all.
The decoder β the fluent one. Never on the request path. Writing and reading memory happen on every turn and stay on the encoder-and-machine side. Generation is for prose a person will read, and for periodic passes that look across many conversations at once to form a persona β a pass that can fail without the conversation stopping.
Where this is heading β one meaning, several machines
Nothing in this section runs in this release. The model you download is the encoder and the Python in this repository; it does not call a VM and contains no circuit. This section says where the work is going, because it explains choices you can already see in the output contract.
The goal is for the reference step itself to have more than one exact implementation. A decision such as "this mention opens a new chain" or "these two mentions are the same entity" is defined once, as a deterministic map on bounded integers, and then carried out by different machines that must agree with that definition:
| machine | what it brings |
|---|---|
| the learned encoder (this model) | reads free text; proposes, per input, which mention points where |
| a deterministic program on a small VM | holds what must never be false β transitivity, one mention in one chain, capacity limits β and is auditable |
| a fixed-point integer circuit compiled from the same program | the same decisions as a neural-style circuit, byte-identical across machines |
The division of labour follows one rule that we measured on this model's own task: invariants go to the exact machine, input-dependent judgements stay with the learned part. Enforcing chain invariants on top of the link scores raised the benchmark score, while the same number of random merges lowered it; turning a per-input judgement into a hand-set constant only moved errors between types once it was trained in.
On the circuit side, small closed controllers already exist in both forms β four lowered by hand and one family lowered automatically by a compiler. For each one checked so far, the integer circuit produces the same output as the program on every reachable state with zero mismatches, and execution can be switched between the program and the circuit at a tick boundary without changing behaviour. That work is on game-agent controllers, not yet on reference resolution.
Why it is not in serving yet. Each machine added to the request path is another thing that has to agree with the others under load. The order is therefore staged:
- stabilise the learned model and the output contract (this release);
- move the chain invariants into a deterministic layer shipped with the model, with no VM;
- run the same invariants as a program on the VM, checked against step 2 for identical output;
- compile that program to an integer circuit and switch between the two at tick boundaries.
A step is taken only after the previous one is measured on the same data, and each step will be reported here when it ships.
Compared with another system
We ran the same public development sets through CorPipe 26
(ufal/corpipe26-onestage-corefud1.4-base-260702, CC BY-NC-SA 4.0 β used only as a point of
comparison, never as a training signal) and scored both sides on linking: for each gold mention,
is the chosen antecedent the previous mention of the same entity, or "none" when it opens a chain?
The data is CorefUD 1.4.
The two sides do not see the same input, so there are two rows per language.
| mentions | this model | CorPipe 26 | gap (pt) | |
|---|---|---|---|---|
| Korean (ko_ecmt) Β· all gold mentions β this model is given the mentions, CorPipe finds its own | 10,263 | 87.8% | 68.5% | +19.4 |
| Korean (ko_ecmt) Β· β only mentions CorPipe also found β linking ability alone | 8,176 | 87.9% | 85.9% | +1.9 |
| English (en_gum) Β· all gold mentions β this model is given the mentions, CorPipe finds its own | 4,641 | 84.6% | 66.0% | +18.5 |
| English (en_gum) Β· β only mentions CorPipe also found β linking ability alone | 3,805 | 84.5% | 80.5% | +4.0 |
- First row β the system as deployed. Mentions reach this model from an upstream component, so it is scored on every gold mention. CorPipe has to find the mentions itself; when it misses one, that row counts it wrong. Most of the gap in this row comes from that difference, not from linking.
- Second row β like for like. Only the mentions CorPipe also found, so both sides are judged on choosing the antecedent. This is the comparison of linking ability.
- It is still not perfectly equal: CorPipe is marked wrong when it found the mention but missed the antecedent mention. Running CorPipe with gold mentions as input would close that, and we have not done it.
What you can check yourself
| 0.2.6 | reproducible here? | |
|---|---|---|
| Japanese reference resolution β | 1.0000 | β |
| CorefUD, 19 languages, macro end-to-end | 0.2240 | β public data |
| input positions this bundle uses | 768 | β config.max_len |
β JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words segmented with janome, 113 in-window cases, mentions supplied. Read this row narrowly: every case in this set has an antecedent and only one candidate, so a model that always links also scores 1.0000 here. It measures whether the model links at all on Japanese input, not whether it chooses well. Rows marked β you can check yourself with the two scripts in this repository.
New knobs (all optional, all defaulting to the old behaviour)
| field | default | what it does |
|---|---|---|
detect_layers |
1 | BIO layers in the detection head. 1 emits outermost mentions only, as before. |
word_joiner |
" " |
What goes between words when the surface is rebuilt. Pass "" for languages written without spaces (Japanese, Chinese, Thai) β otherwise the encoder is handed a surface it has never seen. |
null_bias |
1.0 | How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. Raise it only on a path where every mention really has an antecedent β see the warning below. |
null_bias_by_kind |
{"PRON": 1.0, "DEMNP": 1.0, "NOM": -1.5} |
Per-mention override of null_bias by mention kind (PRON Β· DEMNP Β· NOM, from mention_kind.py). Remove it to get one global value. Passing null_bias= to a call overrides both. |
| Japanese reference resolution | score |
|---|---|
word_joiner="" (as written) |
1.0000 |
word_joiner=" " (words spaced out) |
0.4779 |
Same weights, same inputs, same knob positions otherwise β far apart. A bigger encoder would not have saved you from this.
Each can also be passed per call: model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0).
β Do not set
fix_mistral_regex=TrueRecent
transformersprints a warning when loading this tokenizer, suggesting the flag. We measured it on 5 sentences (ko, ja, en): with the flag, every one of them tokenizes to byte garbage βκ°λ¨μbecomesΓͺ Β° Δ· Γ« Δ€ Β¨instead ofβκ° λ¨ μ, and sequence length goes from 19 to 71. The model still runs and still returns spans; they are simply wrong. The default path is the correct one, and it is the one every number on this card was measured with.
β
null_bias: the number that looks free on the wrong test setOn a benchmark that only ever asks "which earlier mention does this pronoun refer to?", every abstention is wrong by construction, so pushing the model to always answer looks strictly better β and keeps looking better the harder you push. On text where "no antecedent" is sometimes the correct answer, the same push is destructive. We measured both:
null_biasdialogue (KLUE-WoS), mentions given β correct / silent CorefUD ko, linking with gold mentions 3 0.606 / 6 0.745 2 0.599 / 11 0.803 1 (global default) 0.564 / 28 0.844 0.5 0.544 / 48 0.858 0 0.478 / 93 0.871 -0.5 0.379 / 154 0.878 -1 0.269 / 225 0.880 -1.5 0.200 / 286 0.881 by kind β pronoun/demonstrative 1 Β· other noun phrases -1.5 (shipped) 0.564 / 28 0.879 This version ships
null_biasby mention kind: pronouns and demonstratives 1, every other mention -1.5 (the kind is read from the mention's own words, seemention_kind.py). That keeps dialogue at the global-default numbers while full noun phrases stop linking where they should open a new chain. A lower value wins on CorefUD, where many mentions correctly open a new chain and silence is cheap β but on dialogue, where every pointing word has an antecedent, the same setting turns most answers into silence. Accuracy among answered cases barely moves across the range; what moves is how often the model answers at all. If your pipeline guarantees that the mentions you pass in are anaphoric, a higher value is safe; otherwise leave it alone.
License & access
Released under the Elda Community License 1.0 (see LICENSE). Provenance of the training data
is recorded in LINEAGE.md; no user conversations were used.
- β Research, evaluation, internal validation, education β free of charge
- β³ Attribution: "Built with Elda"
- β Commercial use and redistribution require a separate agreement
Access is gated: tell us who you are and access is granted automatically.
Citation
@software{memoryresoner2026,
title = {memory-resoner: multilingual reference resolution for conversational memory},
author = {Elda AI},
year = {2026},
url = {https://huggingface.co/Elda-AI/memory-resoner}
}
- Downloads last month
- 4
# Gated model: Login with a HF token with gated access permission hf auth login